Standard Debug Spec Milestone
Below are milestones planned in the project:
- Phase 1A JTAG DTM + DMI Bus
- Phase 1B Core DM (dmcontrol/dmstatus)
- Phase 1C Debug CSRs
- Phase 1D Triggers
- Phase 2A SWD protocol state machine (
SwdPhy) - Phase 2B SW-DP register file (
SwdDp) - Phase 2C DMI gateway + clock crossing
- Phase 3 Integration (LiteX sim + Arty, OpenOCD, GDB)
- Phase 4 Testing
- Phase 5 Documentation
Status snapshot
| Area | Status |
|---|---|
| Phase 2A–2C (target-side SWD DTM) | ✅ Implemented and sim-verified (sbt + LiteX) |
| Phase 3 step 1 (LiteX + raw-AP smoke) | ✅ DPIDR + dmstatus over SWD |
Phase 3 step 2 (riscv + GDB) | ✅ Vexriscv fork — disdi/openocd vexriscv-gateway (master + Gerrit 9786 + VexRiscv gateway) |
Phase 3 step 2b (demo break / continue) | ✅ Preload path on the Vexriscv fork; no GDB load in sim |
| Phase 3 step 3 (JTAG on Arty) | ✅ 2026-08-20 — tunneled DTM via BSCANE2 USER4; on-board FT2232. Operator how-to: Phase 3 |
| Phase 3 step 3 (SWD on Arty) | ✅ 2026-08-29 — MCU-Link CMSIS-DAP → Pmod JB; smoke + riscv examine + GDB load / break / stepi. Operator how-to: Phase 3 |
| Phase 4–5 | Partial — sim and Arty attach done; full regression matrix and placeholder DPIDR / AP_IDR still open |
Conformance note: the RISC-V Debug Specification defines only a JTAG DTM. SWD is a custom DTM (explicitly permitted). Claim "RISC-V Debug Specification, with custom DTM" — never an unqualified conformance claim.
Phase 1 adds missing features in DebugModule from the ratified RISC-V Debug Specification.
Phase 2 adds a target-side Arm Serial Wire Debug (SWD) front-end that still drives the same RISC-V DebugBus used by the JTAG DTM. The Arm specification is at ADIv5.0–ADIv5.2 (IHI0031).
Host (SWCLK / SWDIO)
│
┌────▼────┐
│ SwdPhy │ Phase 2A — wire protocol (framing, turnaround, line reset)
└────┬────┘
│ SwdDpCmd / SwdDpRsp / SwdDpWrite
┌────▼────┐
│ SwdDp │ Phase 2B — DP registers + ACK policy (OK / WAIT / FAULT)
└────┬────┘
│ SwdApCmd / SwdApRsp
┌────▼──────────┐
│ SwdDmiGateway │ Phase 2C — custom DMI AP + SWCLK↔debug CDC
└────┬──────────┘
│ DebugBus (same as JTAG DTM)
┌────▼────────┐
│ DebugModule │ Phases 1B–1D
└─────────────┘
How the SWD protocol works
What SWD is
Serial Wire Debug is a 2-wire, synchronous, packet-based host↔target link:
| Signal | Role |
|---|---|
| SWCLK | Clock (usually driven by the probe/host) |
| SWDIO | Bidirectional data (one bit per clock) |
There is no separate reset pin. Recovery uses a line reset on SWDIO.
Logically, SWD talks to a Debug Port (DP). The DP either answers for itself (DP registers) or forwards accesses to an Access Port (AP) that reaches the real debug resource. In this project the AP is a custom RISC-V DMI gateway, which drives the existing DebugBus into the Debug Module.
One transaction on the wire
Every SWD access is request → ACK → (optional data).
Packet request (host → target, 8 bits)
After optional idle cycles (SWDIO low), the host sends:
| Bit(s) | Name | Meaning |
|---|---|---|
| 1 | Start | Always 1 |
| 1 | APnDP | 0 = DP register, 1 = AP register |
| 1 | RnW | 0 = write, 1 = read |
| 2 | A[2:3] | Register address bits A[3:2] (LSB-first on the wire) |
| 1 | Parity | Even parity over APnDP, RnW, A[2], A[3] |
| 1 | Stop | Always 0 |
| 1 | Park | Host drives 1, then releases the line |
All multi-bit fields are LSB first.
Turnaround (Trn)
When drive ownership changes, neither side should drive for a short period (default 1 SWCLK). That prevents bus fight on SWDIO.
ACK (target → host, 3 bits)
| Encoding (value) | Name | Wire order (LSB first) |
|---|---|---|
0b001 | OK | 1, 0, 0 |
0b010 | WAIT | 0, 1, 0 |
0b100 | FAULT | 0, 0, 1 |
The numeric value and the bit order on the wire are easy to confuse: OK is value 001, so the first driven bit is 1.
Data phase (only if ACK = OK in this implementation)
| Direction | When | Format |
|---|---|---|
| Write | After OK + second turnaround | 32-bit WDATA + even parity (host → target) |
| Read | After OK, no turnaround (target keeps the line) | 32-bit RDATA + even parity (target → host) |
Errors and special sequences
| Condition | Target behaviour |
|---|---|
| Protocol error (bad request parity / Stop / Park) | Do not drive ACK; stay silent until line reset |
| WDATA parity fail | ACK already sent; DP sets WDATAERR and drops the write (not a line silence) |
| WAIT | Busy (e.g. previous AP access outstanding); host typically retries |
| FAULT | Sticky error set; host clears via ABORT |
| Line reset | SWDIO HIGH for ≥ 50 SWCLK, then ≥ 2 idle (LOW) |
| Idle | SWDIO low between frames, or back-to-back Start with zero idle |
Clock and pins
Clock domain = SWCLK (probe-driven)
swdio.i ← host/probe (or pad input)
swdio.o → pad output value when target drives
swdio.oe → pad output enable
- Target samples
iand updateso/oeon rising SWCLK (OpenOCD bitbang model).
Minimal mental model (one frame)
- Host clocks an 8-bit header (Start…Park).
- Phy validates it; if bad → silence until line reset.
- If good → phy fires cmd and needs rsp.ack for the next few clocks.
- Phy drives ACK (value chosen by Phase 2B).
- If OK and read → phy streams rdata + parity from the DP.
- If OK and write → phy releases the line, samples wdata + parity, fires wr.
- If WAIT/FAULT → phy stops after ACK (no data phase here).
- Host may retry, clear sticky via ABORT or line-reset.
Debug Transport Module
Phase 1a: JTAG System Integration
Spec Feature Coverage Summary
| Spec Feature (Chapter) | Vexriscv Debug Implementation | Implementation Details |
|---|---|---|
| Ch 6: Debug Transport Module (DTM) | ||
| JTAG TAP with IDCODE/dtmcs/dmi | ✅ Implemented | DebugTransportModuleJtag.scala — standard IR codes, dtmcs with version/abits/idle/dmistat, dmi with op/address/data |
| DMI bus protocol | ✅ Implemented | DebugInterfaces.scala — DebugBus with DebugCmd/DebugRsp, DebugBusSlaveFactory |
| DMI busy/error handling | ✅ Implemented | dmihardreset, dmireset, pending/overrun detection |
| Cross-clock-domain DMI | ✅ Implemented | ccToggle for JTAG↔debug clock domains |
| JTAG tunnel support | ✅ Implemented | JtagTunnel.scala — tunneling through outer TAP |
Debug Module
Phase 1B: Core DM Registers
| Spec Feature (Chapter) | Vexriscv Debug Implementation | Implementation Details |
|---|---|---|
| Ch 3: Debug Module (DM) | ||
dmcontrol (0x10) | ✅ Implemented | dmactive, ndmreset, haltreq, resumereq, ackhavereset, hartsello/hi. Missing: hasel, setresethaltreq, hartreset, keepalive |
dmstatus (0x11) | ✅ Implemented | version, authenticated=1, all halted/running/unavail/nonexistent/resumeack/havereset flags, impebreak=1 |
hartinfo (0x12) | ✅ Implemented | dataaddr=0, datasize=0, dataaccess=0, nscratch=0 |
abstractcs (0x16) | ✅ Implemented | datacount, progbufsize, busy, cmderr (all 7 error codes) |
command (0x17) — Access Register | ✅ Implemented | cmdtype=0 with full FSM: transfer, write, postexec, aarsize validation, GPR+FPU register access |
command (0x17) — Access Memory | ❌ Not implemented | Returns NOT_SUPPORTED |
command (0x17) — Quick Access | ❌ Not implemented | Returns NOT_SUPPORTED |
abstractauto (0x18) | ✅ Implemented | autoexecdata and autoexecProgbuf for burst access |
progbuf0-N (0x20+) | ✅ Implemented | Parameterized progbuf memory, multi-word execution with counter, redo support |
data0-N (0x04+) | ✅ Implemented | Memory-backed, hart writes via fromHarts, host reads async |
sbcs (0x38) — System Bus Access | ✅ Implemented (optional) | sbversion=1, sbaccess, sbbusyerror, sbbusy, sbreadonaddr, sbautoincrement, sbreadondata, sberror, 32-bit only |
sbaddress0 (0x39) | ✅ Implemented | Read/write with auto-increment |
sbdata0 (0x3c) | ✅ Implemented | Read/write with bus triggers |
sbaddress1-3 / sbdata1-3 | ❌ Not implemented | 32-bit address/data only |
sbcs 8/16/64/128-bit access | ❌ Not implemented | Only sbaccess32 supported |
haltsum0 (0x40) | ✅ Implemented | Per-hart halted bits, up to 32 harts |
haltsum1-3 | ❌ Not implemented | Only haltsum0 exists |
authdata (0x30) | ❌ Not implemented | Always authenticated |
confstrptr0-3 | ❌ Not implemented | — |
nextdm (0x1d) | ❌ Not implemented | — |
dmcs2 (0x32) | ❌ Not implemented | — |
Hart arrays (hawindowsel/hawindow) | ❌ Not implemented | — |
custom0-15 | ❌ Not implemented | — |
| Multi-hart support | ✅ Implemented | Parameterized p.harts, per-hart buses, hartSel selection |
CSR Register
Phase 1C: Debug CSRs (dcsr, dpc)
| Spec Feature (Chapter) | Vexriscv Debug Implementation | Implementation Details |
|---|---|---|
| Ch 4: Core Debug (hart-side CSRs) | ||
| Halt/Resume | ✅ DebugHartBus + CsrPlugin | CsrPlugin.scala (line 702+): running flag, DebugHartBus wiring, halt/resume handshake |
| Single-step | ✅ Implemented | CsrPlugin.scala (lines 807-845): dcsr.step with full FSM (IDLE→SINGLE→WAIT), timeout/redo handling |
dcsr (0x7B0) | ✅ Implemented | CsrPlugin.scala (lines 792-851): prv, step, nmip, mprven, cause, stoptime, stopcount, stepie, ebreakm/s/u, xdebugver=4 |
dpc (0x7B1) | ✅ Implemented | CsrPlugin.scala (line 791): Reg(UInt(32 bits)), read/write via rw(CSR.DPC, dpc) |
dscratch0 (0x7B2) | ❌ Not implemented | — |
dscratch1 (0x7B3) | ❌ Not implemented | — |
| Debug mode entry | ✅ Implemented | CsrPlugin.scala (lines 1426-1443): saves PC→dpc, sets dcsr.cause (1=ebreak, 3=haltreq, 4=step), saves privilege→dcsr.prv, enters M-mode |
| Debug mode exit (resume) | ✅ Implemented | CsrPlugin.scala (lines 1488-1498): jumps to dpc, restores privilege from dcsr.prv, via DebugHartBus.resume |
| Halt cause reporting | ✅ Implemented | dcsr.cause: 1 (ebreak), 2 (trigger), 3 (haltreq), 4 (step) |
dcsr.ebreakm/s/u | ✅ Implemented | CsrPlugin.scala (lines 1372-1377): per-privilege ebreak→debug detection |
dcsr.stoptime | ✅ Implemented | CsrPlugin.scala (line 866): stoptime output gated by debugMode |
dcsr.stopcount | ✅ Implemented | CsrPlugin.scala (line 1176): mcycle increment gated by !debugMode || !stopcount |
dcsr.stepie | ✅ Implemented | CsrPlugin.scala (line 1315): interrupts cleared when step && !stepie |
| Interrupt inhibition in debug | ✅ Implemented | CsrPlugin.scala (line 721): inhibateInterrupts() when debugMode |
| CSR access protection (0x7Bx) | ✅ Implemented | CsrPlugin.scala (line 1718): blocks non-debug access to 0x7B0-0x7BF |
DebugHartBus wiring | ✅ Implemented | CsrPlugin.scala (lines 704-789): instruction injection, data CSR, all hartToDm/dmToHart signals |
| Reset control | ✅ dmcontrol.ndmreset → io.ndmreset | Both implementations provide ndmreset |
Triggers
Phase 1D: Trigger Module
| Spec Feature (Chapter) | Vexriscv Debug Implementation | Implementation Details |
|---|---|---|
| Ch 5: Trigger Module (hart-side CSRs) | ||
tselect (0x7A0) | ✅ Implemented | CsrPlugin.scala (lines 870-875): WARL index, parameterized debugTriggers (default 2) |
tinfo (0x7A4) | ✅ Implemented | CsrPlugin.scala (line 878): reports type 2 (mcontrol) support |
tdata1 (0x7A1) | ⚠️ Partial | CsrPlugin.scala (lines 922-942): type=2, dmode, execute, m/s/u, action. Missing: timing, select, sizelo/hi, maskmax, chain, match, load, store, hit |
tdata2 (0x7A2) | ✅ Implemented | CsrPlugin.scala (lines 944-953): 32-bit compare value, PC equality match |
tdata3 (0x7A3) | ❌ Not implemented | — |
tcontrol (0x7A5) | ❌ Not implemented | No mte/mpte |
| Trigger type | ⚠️ mcontrol (type 2) only | Legacy type 2, not type 6 (mcontrol6). Spec recommends type 6 for new implementations |
| Match modes | ⚠️ Equal only | Only match=0 (equality). No napot, >=, <, mask modes |
| Match targets | ⚠️ Execute address only | Only execute bit implemented. No load/store data/address match |
| Privilege filtering | ✅ Implemented | m, s, u bits with privilegeHit logic (lines 930-934) |
dmode security | ✅ Implemented | dmode bit controls debug-only write access (line 925) |
action field | ⚠️ Partial | Register exists but only action=1 (enter debug) used in match logic |
| Trigger chaining | ❌ Not implemented | No chain bit |
dcsr.cause=2 on trigger | ✅ Implemented | CsrPlugin.scala (line 892): sets dcsr.cause := 2 on trigger hit |
| Trigger hit → debug entry | ✅ Implemented | CsrPlugin.scala (lines 881-897): decodeBreak halts pipeline, enters debug mode |
mcontrol6 (type 6) | ❌ Not implemented | Only legacy type 2 exists |
icount (type 3) | ❌ Not implemented | — |
itrigger/etrigger/tmexttrigger | ❌ Not implemented | — |
| Hardware breakpoints | ⚠️ Spec-compliant but limited | Type 2 mcontrol with execute address match=0 only, privilege filtering, dmode security |
mcontext/scontext | ❌ Not implemented | — |
| SoC Integration | ||
| DTM→DM→Hart wiring | ✅ Implemented | DebugModuleFiber.scala — multi-hart binding, clock-domain-safe pipelining |
| Tilelink SBA bridge | ✅ Implemented | makeSysbusTilelink() in DebugModuleFiber |
Phase 2A — SWD Protocol State Machine: Implementation & Verification
Code Repository :
Update submodules in pythondata-cpu-vexriscv_smp to below :
- SpinalHdl - https://github.com/disdi/SpinalHDL/tree/phase2a
- VexRiscv - https://github.com/disdi/VexRiscv/tree/phase2a
| Artifact | Path |
|---|---|
| RTL | EXT/SpinalHDL/lib/src/main/scala/spinal/lib/cpu/riscv/debug/DebugTransportModuleSwd.scala |
| Testbench | EXT/VexRiscv/src/test/scala/vexriscv/DebugSwdTest.scala |
| Run | cd EXT/VexRiscv && sbt -batch "testOnly vexriscv.DebugSwdTest" |
EXT = pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext. EXT/VexRiscv/build.sbt compiles EXT/SpinalHDL from source, so the RTL and
testbench build in one sbt project with no LiteX involvement.
1. What is implemented — SwdPhy
SwdPhy is the ARM ADI SW-DP line layer (ADIv6.0 §B4): it speaks the 2-wire protocol
and terminates in a decoded-transaction seam. It contains no DP registers (Phase 2B) and no
DebugBus bridge (Phase 2C).
1.1 I/O
swdio.i : in Bool -- SWDIO as driven by the probe
swdio.o : out Bool -- SWDIO value when the target drives
swdio.oe : out Bool -- target output enable
dp.cmd : master Flow(SwdDpCmd) -- decoded request (2A -> 2B)
dp.rsp : slave Flow(SwdDpRsp) -- ACK + read data (2B -> 2A)
dp.wr : master Flow(SwdDpWrite) -- write commit (2A -> 2B)
- Clock = SWCLK. The component's implicit clock domain is the probe-driven SWCLK.
The target samples
swdio.iand updatesswdio.o/swdio.oeon the rising edge, matching the OpenOCD bitbang model (host sets data while SWCLK is low, samples target data while low). - No
inout. The tristate split (i/o/oe) is required because the cluster is a Verilog black box in LiteX.
1.2 The 2A↔2B seam (three flows)
A wire-protocol fact shapes the seam: a write's ACK is sent before the 33-bit WDATA phase, so write data cannot ride in the request.
| Flow | Fired | Payload |
|---|---|---|
SwdDpCmd | one-cycle pulse on the packet-request park edge, iff parity/stop/park all pass | apNdp, rnw, addr = A[3:2] |
SwdDpRsp | must be presented by the DP within the turnaround cycle (an always-ready DP may simply hold it valid) | ack (OK=001/WAIT=010/FAULT=100), rdata |
SwdDpWrite | one-cycle pulse after the 33rd WDATA bit | data, parityOk (false ⇒ WDATAERR material for 2B) |
The response is latched into a hold register on first sight (rsp.valid may be
combinational off cmd or held continuously); ACK and RDATA are driven from the
latched copy for the rest of the frame.
Read path — READ_DATA → RELEASE → IDLE (Fig B4-2)
- 32 data bits LSB-first from
rspHold.rdata, then even parity (xorR). - No turnaround between ACK and RDATA (target keeps
oe=1). RELEASE:oDrive := False, thenIDLE(trailing Trn / release).
Write path — WR_TRN → WRITE_DATA → IDLE (Fig B4-1)
WR_TRN: two cycles withoe=0(cnt0 then 1) = second turnaround.WRITE_DATA: shift in 32 bits + sample parity; firedp.wrwithdataandparityOk.- WDATA parity fail →
parityOk = false(WDATAERR material for 2B), not protocol error. - Return direct
WRITE_DATA → IDLE(noRELEASE): host already owns the line. Diagram §1.4.2 also usesWRITE_DATA --> IDLE.
ERROR and line reset
| Diagram §1.4.2 | RTL |
|---|---|
ERROR --> ERROR (ignore traffic) | ERROR: oDrive := False only; no header parse |
ERROR --> IDLE on line reset | lineReset.hit → RESET_WAIT → first low → IDLE |
Line reset is orthogonal (overrides any state): ≥50 consecutive highs on swdio.i while !oDrive; counter frozen/cleared while target drives so ACK/RDATA cannot fake a reset.
RESET_WAIT is an RTL-only gate so the target does not accept Start until the line has gone idle after the reset burst. Diagram §1.4.2 draws a direct ERROR → IDLE; recovery contract is the same.
2. How the testbench works — DebugSwdTest
The bench is the probe
SWCLK is the DUT clock, and the bench owns it: no forkStimulus — every SWCLK cycle is
one call to step(bit):
fallingEdge(); set swdio.i = bit; // host updates while SWCLK low
sample (swdio.o, swdio.oe); // host samples while SWCLK low
risingEdge(); // target samples/updates
This reproduces OpenOCD's bitbang_swd_exchange exactly: a host-driven bit is sampled by
the target at the rising edge ending its cycle; a target-driven bit read in cycle k is
the value the target registered at edge k−1. All multi-bit values are sent/collected
LSB first. step returns (o, oe), so every helper can assert drive/release behavior
per cycle.
Phase 2B — SW-DP Register File: Implementation & Verification
Code Repository :
Update submodules in pythondata-cpu-vexriscv_smp to below :
- SpinalHdl - https://github.com/disdi/SpinalHDL/tree/phase2b
- VexRiscv - https://github.com/disdi/VexRiscv/tree/phase2b
| Artifact | Path |
|---|---|
RTL (SwdDp, SwdPhyDp, AP seam bundles) | EXT/SpinalHDL/lib/src/main/scala/spinal/lib/cpu/riscv/debug/DebugTransportModuleSwd.scala (same file as Phase 2A) |
| Shared probe driver | EXT/VexRiscv/src/test/scala/vexriscv/SwdSimDriver.scala |
| Testbench | EXT/VexRiscv/src/test/scala/vexriscv/DebugSwdDpTest.scala |
| Run | cd EXT/VexRiscv && sbt -batch "testOnly vexriscv.DebugSwdDpTest" (both phases: "testOnly vexriscv.DebugSwdTest vexriscv.DebugSwdDpTest") |
EXT = pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext. Still zero
LiteX involvement.
1. What is implemented
1.1 SwdDp — the ADI SW-DP register file
SwdDp sits behind the Phase 2A seam (SwdDpCmd/SwdDpRsp/SwdDpWrite, consumed as
slave flows) and exposes a new AP seam toward Phase 2C:
ap.cmd : master Flow(SwdApCmd) -- rnw, addr = A[3:2], apSel = SELECT[31:24], wdata
ap.rsp : slave Flow(SwdApRsp) -- error, data (completion; test-controlled latency)
1.2 DP register map (SWD A[3:2] encoding)
| A[3:2] | Read | Write |
|---|---|---|
00 | DPIDR (parameter; default 0x0BA11AAB placeholder) | ABORT — DAPABORT, STKCMPCLR, STKERRCLR, WDERRCLR, ORUNERRCLR |
01 | CTRL/STAT when DPBANKSEL==0; other banks read as zero | CTRL/STAT when DPBANKSEL==0; other banks write-ignored |
10 | RESEND (= last posted result) | SELECT — APSEL[31:24], APBANKSEL[7:4], DPBANKSEL[3:0] |
11 | RDBUFF (= last posted result) | TARGETSEL (SWD v2) — accepted, ignored |
These two diagrams put it back on the
wire, making explicit where each step of a DP access is accomplished: the frame shape is
Phase 2A (SwdPhy), the register semantics are Phase 2B (SwdDp). Neither phase performs a
DP access on its own.
A DP read — DPIDR, A[3:2] = 00
Wire layout — one cell per SWCLK bit period, LSB first within every multi-bit field:
Bits 0–7 host-driven · bits 9–44 target-driven · bits 8 and 45 are turnarounds where
neither side drives and SWDIO floats to its mandatory pull-up. The closing turnaround is
one bit period on the wire but two SWCLK cycles in RTL (RELEASE), because SwdPhy
registers its outputs — see the step-ownership table below.
Phase ownership:
A DP write — SELECT, A[3:2] = 10
Wire layout — note the second turnaround at bit 12, absent from the read:
Bits 0–7 host-driven · bits 9–11 target-driven · bits 13–45 host-driven again · bits 8 and 12 are turnarounds. This layout is the reason the 2A↔2B seam has three flows: the target must commit to the ACK at bits 9–11, thirty-two bit periods before the data it is acknowledging exists on the wire.
Phase ownership:
2. Testbench changes
2.1 Shared driver extraction
The probe-side bit-bang logic (step, header, readAck, transactRead,
transactWrite, lineReset, with all embedded protocol assertions) moved from the 1a
suite into SwdSimDriver.scala (SwdHostDriver, SwdAckSim). DebugSwdTest
delegates to it with its 9 test bodies unchanged — proven by the suites running together
(18/18). Step 2 will reuse the same driver against the full transport + DebugModule.
2.2 The stub moves back one layer
Exactly as the step-1b plan prescribes: the always-ready DP stub is gone — the real
SwdDp answers within the turnaround by construction (combinational response). The test
harness now stubs the AP side:
- an observer thread records every
ap.cmdfire into a queue ((rnw, addr, wdata)); apComplete(data, error)delivers a completion onap.rspfor exactly one SWCLK cycle — completions are test-controlled, which is what makes posted-read ordering and WAIT-while-busy directly testable (the AP simply doesn't respond until the test says so).
Sugar wrappers keep tests readable: dpRead/dpWrite/apRead/apWrite map to driver
transactions with APnDP set accordingly.
4. Verification
cd ~/fpga/pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext/VexRiscv
sbt -batch "testOnly vexriscv.DebugSwdTest vexriscv.DebugSwdDpTest"
# expect: Tests: succeeded 18, failed 0, canceled 0, ignored 0, pending 0
Phase 2C — DMI Gateway, Clock Crossing & DebugModule Integration
Code Repository :
- pythondata-cpu-vexriscv_smp - https://github.com/disdi/pythondata-cpu-vexriscv_smp/tree/phase2c
OR
Update submodules in pythondata-cpu-vexriscv_smp to below :
- SpinalHdl - https://github.com/disdi/SpinalHDL/tree/phase2c
- VexRiscv - https://github.com/disdi/VexRiscv/tree/phase2c
| Artifact | Path |
|---|---|
RTL (SwdDmiGateway, DebugTransportModuleSwd) | EXT/SpinalHDL/lib/src/main/scala/spinal/lib/cpu/riscv/debug/DebugTransportModuleSwd.scala (same file as 2A/2B) |
Fiber hook (withSwdTransport()) | EXT/SpinalHDL/lib/src/main/scala/spinal/lib/cpu/riscv/debug/DebugModuleFiber.scala |
Testbench (incl. SwdDmTestTop) | EXT/VexRiscv/src/test/scala/vexriscv/DebugSwdDmTest.scala |
| Run | cd EXT/VexRiscv && sbt -batch "testOnly vexriscv.DebugSwdDmTest" (all: "testOnly vexriscv.DebugSwdTest vexriscv.DebugSwdDpTest vexriscv.DebugSwdDmTest") |
EXT = pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext. Still zero LiteX involvement.
1.Implementation of DMI for SWD
1.1 SwdDmiGateway — the RISC-V DMI gateway AP
Sits behind the 2B AP seam (SwdApCmd/SwdApRsp) and drives a DebugBus master —
the same DebugCmd/DebugRsp interface the JTAG DTM produces, which is the whole
point of the transport-agnostic design.
The DMI needs a 7-bit address and 32 bits of data. SWD offers, per transaction, one bit of DP/AP selection and two bits of register address. Four AP registers, total.
AP A[3:2] | Register | Access | Behavior |
|---|---|---|---|
00 | AP_IDR | RO | identification constant (default 0x74726976 "triv") — completes locally in 1 SWCLK cycle |
01 | DMI_ADDR | RW | latches the 7-bit DMI word address; readable back — local |
10 | DMI_DATA | RW | performs the DebugBus transaction at DMI_ADDR (RnW from the SWD packet) |
11 | POSTED_READ | RO | last completed DMI read result — local |
Writes to RO registers complete OK with no effect. SELECT.APSEL is not decoded
(single-AP design). AP reads launch at the request (posted, per 2B); AP writes launch at
the WDATA commit carrying the data.
1.2 DebugTransportModuleSwd — the full transport
SWD pins → SwdPhy (2A) → SwdDp (2B) → SwdDmiGateway (2C) → DebugBus. The SWD side
elaborates under ClockDomain(swclk, resetKind = BOOT); the DebugBus side under the
provided debugCd. This is the SWD counterpart of DebugTransportModuleJtagTap.
DebugModuleFiber.withSwdTransport(dpidr, apIdr) instantiates it alongside
withJtagTap() — same one-transport-per-build rule (direct io.ctrl connection, no
arbitration).
1.3 End-to-end dmstatus read (the step-2 exit), graphically
2. Testbench — DebugSwdDmTest
- DUT =
SwdDmTestTop: the full transport + a realDebugModule(version 2, 1 hart,progBufSize=2,datacount=1) with a stubbedDebugHartBus— running, never halted,hartToDm/resume.rspidle. - Two genuinely asynchronous clocks: the shared
SwdHostDriverbit-bangs SWCLK (bench-owned, via aClockDomainhandle over the pin) whiledut.clockDomainfree-runs viaforkStimulus— DMI completion latency is variable by construction. - Host-style retry helpers:
apWriteRetry/apReadRetry/rdbuffretry on WAIT with bounded attempts — the same loop a real OpenOCD target would run.dmiRead/dmiWritecompose them into DMI operations. - Failure forensics built in: on any unexpected ACK the harness reads CTRL/STAT and
includes the sticky flags in the assertion message (
ackCheck) — this is what cracked the phantom-completion bug (§4.2). Thesim(name, seed = …)hook pins a failing seed for deterministic reproduction.
3. Verification
cd ~/fpga/pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext/VexRiscv
sbt -batch "testOnly vexriscv.DebugSwdDmTest"
# expect: Tests: succeeded 7, failed 0, canceled 0, ignored 0, pending 0
sbt -batch "testOnly vexriscv.DebugSwdTest vexriscv.DebugSwdDpTest vexriscv.DebugSwdDmTest"
# expect: Tests: succeeded 25, failed 0, canceled 0, ignored 0, pending 0
4. Summary
Phases 2A–2C complete the target-side RTL of the SWD plan; the DebugBus now speaks SWD end-to-end in simulation.
System Integration — SWD
Goal is host-visible integration only (LiteX SoC, OpenOCD, GDB). Target-side SWD
RTL (SwdPhy / SwdDp / SwdDmiGateway) is Phase 2A–2C and is not re-defined here.
Code repositories
| Piece | Where | Status |
|---|---|---|
SWD DTM (DebugTransportModuleSwd) + DebugModuleFiber.withSwdTransport() | SpinalHDL #1956 | merged 2026-09-08 |
spinal.lib.com.swd split (Swd / SwdPhy / SwdDp) | SpinalHDL #1966 | merged 2026-09-19 |
VexRiscv SMP cluster --swd | VexRiscv #483, #499 | merged 2026-09-08 / 09-19 |
VexiiRiscv LiteX SoC --with-swd + MicroSoc --swd | VexiiRiscv #184 | merged 2026-09-23 |
VexiiRiscv CPU-embedded debug plugin (EmbeddedRiscvJtag) over SWD, --debug-swd | VexiiRiscv #188 | merged 2026-09-28 (19b41a7) |
VexRiscv CPU-embedded debug plugin (EmbeddedRiscvJtag) over SWD | VexRiscv #500 | merged 2026-09-27 (aefc0e0) |
ElemRV / nafarr: DebugTransport (Jtag default / Swd) on the VexiiRiscv realtime and performance presets | elements-nafarr #72 | merged 2026-09-29 (54406843). VexiiRiscv pin 19b41a7 is #74 (b7a1257) |
ElemRV / zibal: debugTransport on the Hydrogen / Carbon / Nitrogen platforms → io_plat.jtag or io_plat.swd | elements-zibal #62 | open, submitted 2026-09-29 (d78489b on main 8f5e3f5, which already pins nafarr 54406843) |
JTAG on Xilinx USER chains (add_cpu_jtag_debug, --with-cpu-jtag-debug) | LiteX #2572 + linux-on-litex-vexriscv #459 | merged 2026-09-10 |
LiteX: swdremote sim module, OpenOCD configs, --with-swd-debug for vexriscv_smp | https://github.com/disdi/litex/tree/swd | branch, not yet proposed upstream |
| linux-on-litex-vexriscv: SWD pads on Arty Pmod JB | https://github.com/disdi/linux-on-litex-vexriscv/tree/swd-arty | branch, waits for the LiteX part |
| OpenOCD (Vexriscv fork) | https://github.com/disdi/openocd/tree/vexriscv-gateway | branch; Gerrit 9786 + gateway backend |
| Black Magic Debug: DMI gateway AP support (no OpenOCD) | blackmagic #2322 — disdi:feature/riscv-swd-dmi-gateway | open, submitted 2026-09-25 — see Black Magic Probe |
OpenOCD (host ONLY) — :
| Lane | Build | Role |
|---|---|---|
| raw-AP smoke | OpenOCD master (stock OK for SWD remote_bitbang) | DPIDR + dap apreg → dmstatus; no GDB |
Vexriscv fork (riscv + GDB) | disdi/openocd vexriscv-gateway — OpenOCD master + Gerrit 9786 + designer-AP / VexRiscv DTM backend | examine + halt/resume/regs + GDB :3333 |
Stock master is enough for smoke. Full riscv attach needs the published
vexriscv-gateway branch:
9786 (DTM + Mem-AP DMI backend), two fixes to 9786 itself, and a second designer-AP /
VexRiscv gateway backend for the Phase 2C DMI_ADDR / DMI_DATA map. No RTL change.
Upstream’s stock riscv target rejects -dap at argument parsing, so Tcl-only
dap apreg helpers cannot drive a GDB session — that is why the 9786 + gateway path exists.
Simulation based workflow using verilator
Side-by-side
| Terminal | JTAG — full three-terminal ✅ | SWD — full three-terminal ✅ |
|---|---|---|
| 1 — sim | litex_sim … --with-privileged-debug --jtag-tap --with-jtagremote → TCP 44853 (jtagremote) | litex_sim … --with-privileged-debug --with-swd-debug --with-swdremote → TCP 44854 (swdremote); add --ram-init=demo.bin for demo debug |
| 2 — OpenOCD | Stock riscv target + fabric TAP — no vendor BSCAN in sim; examines hart; GDB :3333 | transport select swd + DAP + Vexriscv fork riscv; examines hart; GDB :3333 |
| 3 — GDB | target extended-remote localhost:3333 → halt / regs / load | attach + regs ✅; demo break main / continue / bt ✅ via preload (no GDB load) |
| Capability | JTAG | SWD |
|---|---|---|
| Verilator SoC + official DM | ✅ | ✅ (_Swd cluster) |
| Wire transport in sim | ✅ JTAG TAP + tunnel | ✅ SW-DP (DebugTransportModuleSwd) |
| OpenOCD sees transport | ✅ TAP 0x10003fff | ✅ SWD DPIDR 0x0ba11aab |
Read dmstatus | ✅ via riscv / DMI | ✅ via vexriscv_dmi_read 0x11 / smoke |
Examined RISC-V core | ✅ | ✅ XLEN=32, misa=0x40141101 |
GDB halt / resume / info registers | ✅ | ✅ |
| Break / continue / backtrace | ✅ (load OK on JTAG) | ✅ via --ram-init=demo.bin + symbols; GDB load impractical in sim |
| OpenOCD binary | stock master | stock master for raw-AP smoke; Vexriscv fork (disdi/openocd vexriscv-gateway) for riscv / GDB |
JTAG — end-to-end workflow
Official stack only (--with-privileged-debug + full JTAG TAP in sim).
Prerequisites
litex_sim(LiteX venv)- OpenOCD master with standard RISC-V target
riscv64-unknown-elf-gdb
Configs
| File | Role |
|---|---|
openocd_jtag_remote.cfg | remote_bitbang → localhost:44853 |
riscv_jtag_tunneled.tcl | TAP irlen 6, ID 0x10003fff, riscv use_bscan_tunnel 6 1 |
Terminal 1 — sim (keep running)
litex_sim \
--integrated-main-ram-size=0x10000 \
--cpu-type=vexriscv_smp \
--cpu-variant=linux \
--cpu-count=1 \
--with-privileged-debug \
--jtag-tap \
--with-jtagremote \
--non-interactive
| Flag | Role |
|---|---|
--integrated-main-ram-size=0x10000 | 64 KiB main RAM for sim (demo load region) |
--cpu-type=vexriscv_smp / --cpu-variant=linux / --cpu-count=1 | SMP Linux-capable cluster, 1 hart |
--with-privileged-debug | Official DebugModule + DTM (_Pd netlist token) |
--jtag-tap | Full JTAG TAP on cluster (_JtagT); needed so sim has TCK/TMS/TDI/TDO pads |
--with-jtagremote | LiteX sim module jtagremote — OpenOCD remote_bitbang on TCP 44853 |
--non-interactive | Keep sim running (no local control menu); target for OpenOCD/GDB |
Wait for Found port 44853 and BIOS prompt litex>. First run may take several minutes
(cluster regen + Verilator compile).
Terminal 2 — OpenOCD (after Terminal 1 is up)
openocd -f openocd_jtag_remote.cfg -f riscv_jtag_tunneled.tcl
Success indicators:
Info : JTAG tap: riscv.cpu tap/device found: 0x10003fff
Info : Examined RISC-V core; found 1 harts
Ready for Remote Connections
Info : Listening on port 3333 for gdb connections
Terminal 3 — GDB (after OpenOCD is ready)
riscv64-unknown-elf-gdb demo/demo.elf
set remotetimeout 120
set pagination off
set arch riscv:rv32
target extended-remote localhost:3333
monitor reset halt
x/8i $pc
info registers
One-liner:
riscv64-unknown-elf-gdb -ex "set remotetimeout 120" \
-ex "target extended-remote localhost:3333" \
demo/demo.elf
Load demo.elf (linked at 0x40000000) only after halt — this litex_sim invocation does
not pass --ram-init=demo.bin, so the image is not preloaded:
monitor reset halt
load demo/demo.elf
break main
continue
If sim is slow and keep_alive() warnings (slow bitbang) are seen, prefer target extended-remote
and set remotetimeout 120.
SWD — OpenOCD
| Lane | OpenOCD build | Configs | Gives you |
|---|---|---|---|
| raw AP | master | openocd_swd_remote.cfg + vexriscv_swd.cfg | DPIDR + dap apreg → dmstatus; no GDB |
riscv Vexriscv fork | disdi/openocd vexriscv-gateway (master + 9786 + gateway) | + vexriscv_swd_riscv_master.cfg | examine + halt/resume/regs + GDB :3333 |
Prerequisites (SWD-specific)
litex_sim- OpenOCD master with SWD
remote_bitbangfor smoke test - For
riscv/ GDB: build from https://github.com/disdi/openocd/tree/vexriscv-gateway (stock master rejectsriscv -dap) riscv64-unknown-elf-gdb; useset remotetimeout 300on the SWD lane
Configs
| File | Role |
|---|---|
openocd_swd_remote.cfg | remote_bitbang → localhost:44854, transport select swd |
vexriscv_swd.cfg | SW-DP + DAP + vexriscv_dmi_read/write + vexriscv_swd_smoke — no riscv target |
vexriscv_swd_riscv_master.cfg | Vexriscv fork: dtm create -type vexriscv-gateway + riscv + gdb-attach halt |
SWD Reading dmstatus, all the way down
Terminal 1 — sim (keep running)
| Goal | Extra flag |
|---|---|
| Attach / regs / raw-AP smoke | (none) — wait for Found port 44854 + BIOS litex> |
Debug the demo app (break main / continue / bt) | --ram-init=demo.bin — wait for serialboot timeout → Executing booted program at 0x40000000 → litex-demo-app> |
litex_sim \
--integrated-main-ram-size=0x10000 \
--cpu-type=vexriscv_smp \
--cpu-variant=linux \
--cpu-count=1 \
--with-privileged-debug \
--with-swd-debug \
--with-swdremote \
--non-interactive
--ram-init=demo.bin
| Flag | Role |
|---|---|
--with-privileged-debug | Official DebugModule (required; SWD is official-stack only) |
--with-swd-debug | Cluster SWD transport + _Swd netlist token |
--with-swdremote | LiteX sim module swdremote — OpenOCD SWD bitbang on TCP 44854 |
--ram-init=demo.bin | Preload demo into main_ram @ 0x40000000 (demo-debug only) |
Terminal 2 — raw-AP smoke (stock master; no GDB)
One-shot:
openocd -s tcl \
-f openocd_swd_remote.cfg \
-f vexriscv_swd.cfg \
-c init -c vexriscv_swd_smoke -c shutdown
Verified success:
Info : SWD DPIDR 0x0ba11aab
AP_IDR = 0x74726976
dmstatus = 0x004c0c82 (version=2 authenticated=1 allrunning=1 allhalted=0)
PASS: SWD -> SW-DP -> DMI gateway -> DebugModule
Interactive Tcl helpers (same two configs, stay open):
vexriscv_swd_smoke
vexriscv_dmi_read 0x11 ;# dmstatus
Terminal 2 — riscv target + GDB server (Vexriscv fork: disdi/openocd vexriscv-gateway)
# Use the openocd binary built from:
# https://github.com/disdi/openocd/tree/vexriscv-gateway
openocd -s tcl \
-f openocd_swd_remote.cfg \
-f vexriscv_swd.cfg \
-f vexriscv_swd_riscv_master.cfg
Verified examine (real hart, not stub):
Info : SWD DPIDR 0x0ba11aab
Info : [vexriscv.rv] datacount=1 progbufsize=2
Info : [vexriscv.rv] Examined RISC-V core
Info : [vexriscv.rv] XLEN=32, misa=0x40141101
vexriscv.rv halted due to debug-request.
misa=0x40141101 = RV32 I+M+A+S+U. Check halt/resume by curstate, not only by log
lines: resume may print halted due to single-step. while stepping off a
breakpoint — that is not a failure.
Terminal 3 — GDB (after the Vexriscv fork is listening on :3333)
Start GDB with the ELF for symbols (attach-only or demo-debug):
riscv64-unknown-elf-gdb demo/demo.elf
Attach and inspect
set remotetimeout 300
set pagination off
set arch riscv:rv32
target extended-remote localhost:3333
info registers
x/6i $pc
vexriscv_swd_riscv_master.cfg sets -event gdb-attach halt, so GDB attaches to an
already-halted target — no monitor halt is required.
Debug the demo app (break / continue / bt)
If litex_sim is passed with --ram-init=demo.bin :
set remotetimeout 300
set pagination off
set arch riscv:rv32
target extended-remote localhost:3333
# Image is already in main_ram via --ram-init=demo.bin.
x/8xw 0x40000000 # confirm preload (e.g. 0x0b00006f 0x00000013 ...)
set $pc = 0x40000000 # re-enter demo at _start so main is hit cleanly
break main
continue
bt
Expected: stop at main (typically around 0x4000069c); bt shows #0 main ().
Hardware based workflow using Arty
Reference manual: https://digilent.com/reference/programmable-logic/arty-a7/reference-manual
| Device | xc7a35ticsg324-1L (a7-35 variant) |
| System clock | clk100, 100 MHz |
| USB | On-board FTDI FT2232HQ — USB-JTAG (programming) + USB-UART (console) on one cable |
| Programmer | OpenOCD via openocd_xc7_ft2232.cfg + bscan_spi_xc7a35t.bit |
| PMOD connectors | JA/JB/JC/JD → pmoda/pmodb/pmodc/pmodd |
| SWD probe | External CMSIS-DAP (MCU-Link) on Pmod JB — not the on-board FTDI |
JTAG on Arty — end-to-end workflow
Official stack only (--with-privileged-debug).
--with-privileged-debug alone is not enough on hardware. Without --jtag-tap the DTM is
tunneled, and its debugPort_* signals need a vendor boundary-scan primitive (BSCANE2 on a
Xilinx USER chain). That binding is now upstream as an explicit second flag,
--with-cpu-jtag-debug — LiteX #2572
(LiteXSoC.add_cpu_jtag_debug(), default USER4, IR 0x23) and linux-on-litex-vexriscv
#459. USER1 stays free for
jtagbone. It replaces the original
#458, which was closed unmerged;
the hardware results below were taken with #458's BSCANE2 instance, the same USER4 / IR 0x23
binding.
Prerequisites
- Arty
- OpenOCD master with standard RISC-V target
riscv64-unknown-elf-gdb
Configs
| File | Role |
|---|---|
openocd_arty_bscan.cfg | Fits on-board FT2232 into the Xilinx TAP |
openocd_arty_official.cfg | JTAG on Arty hardware. Composes openocd_arty_bscan.cfg (adapter + Xilinx TAP) with riscv_jtag_tunneled.tcl (the riscv target + use_bscan_tunnel). |
Terminal 1 — Build and Flash on Arty
# Stock LiteX + linux-on-litex-vexriscv master (litex#2572 / linux-on-litex-vexriscv#459)
./make.py --board=arty --cpu-count=1 --with-privileged-debug --with-cpu-jtag-debug --build --load
Expected in OpenOCD: Examined RISC-V core; found 1 harts and XLEN=32, misa=0x40141101.
Terminal 2 — OpenOCD (after Terminal 1 is up)
openocd -f litex/litex/tools/debug/openocd_arty_official.cfg
Terminal 3 — GDB (after OpenOCD is ready)
riscv64-unknown-elf-gdb -ex "set arch riscv:rv32" -ex "target extended-remote localhost:3333" linux-on-litex-vexriscv/build/arty/software/bios/bios.elf
SWD on Arty — end-to-end workflow
Official stack only (--with-swd-debug; that flag implies --with-privileged-debug).
One transport per bitstream — do not combine with --jtag-tap.
--with-swd-debug alone is not enough on hardware. Unlike JTAG (which reuses the on-board FT2232 as a tunneled TAP), SWD needs an external CMSIS-DAP probe and two user I/O pins. Pinout lives in
soc_linux.py
(_swd_pmod_io / add_cpu_swd_debug) on
https://github.com/disdi/linux-on-litex-vexriscv/tree/swd-arty.
Prerequisites
- Arty
- CMSIS-DAP probe (MCU-Link is the verified one) + jumper wires
- OpenOCD master for the raw-AP smoke; the Vexriscv fork
(disdi/openocd
vexriscv-gateway) for GDB support. riscv64-unknown-elf-gdb
Hardware connection
Two USB cables to the host. The FTDI does not carry SWD.
Host PC
├─ USB ── FT2232 (Arty J10) ── bitstream load + UART console
└─ USB ── MCU-Link (CMSIS-DAP)
SWCLK ──► JB3 ──► cluster swd_clk ──► SwdPhy
SWDIO ◄─► JB7 ── IOBUF (swdio_i / o / oe)
└── SwdPhy → SwdDp → DMI gateway → DebugModule
--with-swd-debug brings SWCLK / SWDIO out on Pmod JB (high-speed header: no 200 Ω series resistors, which matters for bidirectional SWDIO turnaround). Do not use JA or JD.
| Signal | LiteX pin | Pmod JB | FPGA ball | Notes |
|---|---|---|---|---|
| SWCLK | pmodb:2 | JB3 | D15 | Probe-driven, gated clock |
| SWDIO | pmodb:4 | JB7 | J17 | Bidirectional; FPGA PULLUP TRUE (ADI) |
| GND | — | JB11 | — | Common ground (required) |
| GND (2nd) | — | JB5 | — | Required — second ground return, see below |
| 3.3 V (VTref) | — | JB12 | — | Required for the MCU-Link — see below |
Looking into the 12-pin Pmod:
JB1 JB2 JB3=SWCLK JB4 JB5=GND(2nd) JB6=3V3
JB7=SWDIO JB8 JB9 JB10 JB11=GND JB12=3V3 (VTref)
MCU-Link 10-pin Cortex debug header → Arty:
MCU-Link pin 4 (SWCLK) → Arty JB3
MCU-Link pin 2 (SWDIO) → Arty JB7
MCU-Link pin 3 (GND) → Arty JB11
MCU-Link pin 5 (GND) → Arty JB5 (second ground, required)
MCU-Link pin 1 (VTref) → Arty JB12 (required on the MCU-Link)
Configs
| File | Role |
|---|---|
openocd_arty_swd.cfg | CMSIS-DAP adapter. Hardware sibling of openocd_swd_remote.cfg — same DAP/DMI helpers, remote_bitbang swapped for cmsis-dap + usb_bulk. |
vexriscv_swd.cfg | SW-DP + DAP + vexriscv_dmi_read/write + vexriscv_swd_smoke — reused from sim, unchanged |
vexriscv_swd_riscv_master.cfg | Vexriscv fork: dtm create -type vexriscv-gateway + riscv + gdb-attach halt — reused from sim, unchanged |
Terminal 1 — Build and Flash on Arty
# Support for SWD added to https://github.com/disdi/linux-on-litex-vexriscv/tree/swd-arty
./make.py --board=arty --cpu-count=1 --with-swd-debug --build --load
The first SWD build regenerates the cluster via sbt (no _Swd netlist ships prebuilt).
Expected name: VexRiscvLitexSmpCluster_Cc1_…_Ood_Pd_Hb1_Swd.
Terminal 2 — raw-AP smoke (stock master; no GDB)
openocd -s tcl \
-f litex/litex/tools/debug/openocd_arty_swd.cfg \
-f litex/litex/tools/debug/vexriscv_swd.cfg \
-c init -c vexriscv_swd_smoke -c shutdown
Success indicators:
Info : SWD DPIDR 0x0ba11aab
AP_IDR = 0x74726976
dmstatus = 0x004c0c82 (version=2 authenticated=1 allrunning=1 allhalted=0)
PASS: SWD -> SW-DP -> DMI gateway -> DebugModule
dmstatus is bit-identical to the sim value.
Terminal 2 — riscv target + GDB server (Vexriscv fork: disdi/openocd vexriscv-gateway)
# Use the openocd binary built from:
# https://github.com/disdi/openocd/tree/vexriscv-gateway
openocd -s tcl \
-f litex/litex/tools/debug/openocd_arty_swd.cfg \
-f litex/litex/tools/debug/vexriscv_swd.cfg \
-f litex/litex/tools/debug/vexriscv_swd_riscv_master.cfg
Success indicators:
Info : SWD DPIDR 0x0ba11aab
Info : [vexriscv.rv] Examined RISC-V core
Info : [vexriscv.rv] XLEN=32, misa=0x40141101
misa=0x40141101 = RV32 I+M+A+S+U.
Terminal 3 — GDB (after OpenOCD is ready)
riscv64-unknown-elf-gdb -ex "set arch riscv:rv32" -ex "target extended-remote localhost:3333" linux-on-litex-vexriscv/build/arty/software/bios/bios.elf
Hardware is not limited the way sim is: GDB load and stepi both work (CMSIS-DAP at
1 MHz vs swdremote pacing). Debug demo.elf (linked at 0x40000000):
monitor halt
load demo/demo.elf
set $pc = 0x40000000
break main
continue
bt
Expected: stop at main (typically around 0x4000069c); bt shows #0 main ().
VexiiRiscv over SWD
The same SWD transport and DebugModule now also serve VexiiRiscv
VexiiRiscv#184. No new
transport RTL was needed. VexiiRiscv builds its debug logic from SpinalHDL's
DebugModuleSocFiber, and the SWD DTM is added in that fiber's body with
dm.withSwdTransport(). SwdPhy / SwdDp / SwdPhyDp in the generated netlist are
byte-identical to the VexRiscv SMP cluster's, so everything on the host side (probe, OpenOCD fork,
configs) is reused unchanged. The SWD transport is the same SpinalHDL code in both CPUs, and the generated Verilog confirms it.
| SoC | Option | Top-level ports |
|---|---|---|
LiteX SoC (vexiiriscv.soc.litex.SocGen) | --with-swd | debug_swd_swd_swclk, debug_swd_swd_swdio_{read,write,writeEnable} |
MicroSoc (vexiiriscv.soc.micro.MicroSocGen) | --jtag-tap=false --swd=true | socCtrl_debugModule_swd_swd_* |
- One DTM at a time (RISC-V Debug Spec Ch. 6):
--with-swdtogether with--with-jtag-tap/--with-jtag-instructionis refused at elaboration. - SWDIO is three wires (
read/write/writeEnable); the tristate belongs to the integrator. - A
--with-jtag-tapnetlist is unchanged by #184.
LiteX integration — pending. LiteX's cpu/vexiiriscv still pins a VexiiRiscv revision
without SWD, and its --with-swd-debug option for VexiiRiscv (four ports, add_swd(), reset
wiring) is not published yet; it will be proposed together with the pin bump.
Hardware results
Same MCU-Link, wiring and OpenOCD fork as the VexRiscv lane above.
| Configuration | Board | Build | Result |
|---|---|---|---|
RV32 linux (RV32IMA + S/U), 1 hart | Arty A7-35T, 100 MHz | WNS 0.233 ns, 41.6 % LUT | DPIDR 0x0ba11aab, AP_IDR 0x74726976, misa=0x40141101; halt / step / resume; GDB load (67 KB/s), break main / help, stepi, bt |
RV64 debian (RV64IMAFDC + S/U), 1 hart | Arty A7-35T, 100 MHz | WNS 0.131 ns, 75 % LUT / 86 % BRAM | XLEN=64, misa=0x800000000014112d; halt / step / resume; GDB (set arch riscv:rv64): load, breakpoints, 64-bit register write / read-back, FPU registers (fcsr, ft0), bt |
RV64 debian, 2 harts | Arty A7-100T, 80 MHz | WNS 0.532 ns, 44 % LUT | both harts on one DTM; independent halt / step / resume per hart; GDB sees one thread per hart |
Two harts of the RV64 variant do not fit the A7-35T (estimated ~128 % LUT, ~126 % BRAM), hence A7-100T at 80 MHz is used which meets timing.
Log lines that are expected on VexiiRiscv:
Found 0 triggers— LiteX's default VexiiRiscv configuration has no hardware triggers (same over JTAG); software breakpoints in RAM work.Failed to read memory (addr=0x3ffffffc)— GDB peeks at the word before_start, which is unmapped; the DM correctly reports the bus error.Core N could not be made part of halt group 1(two harts) — this DM implements no halt groups, which the spec allows; OpenOCD halts the harts one after the other.
Two harts: OpenOCD configuration
One riscv target per hart, all on the same DTM (DMI is shared by every hart behind the DM;
the hart is chosen by -coreid, i.e. hartsel). Source it in place of
vexriscv_swd_riscv_master.cfg, after openocd_arty_swd.cfg and vexriscv_swd.cfg.
For GDB, one thread per hart:
dtm create vexriscv.dtm -type vexriscv-gateway -dap vexriscv.dap -ap-num 0
target create vexriscv.rv0 riscv -dtm vexriscv.dtm -coreid 0 -rtos hwthread
target create vexriscv.rv1 riscv -dtm vexriscv.dtm -coreid 1 -rtos hwthread
target smp vexriscv.rv0 vexriscv.rv1
riscv set_command_timeout_sec 120
vexriscv.rv0 configure -event gdb-attach halt
set arch riscv:rv64
target extended-remote localhost:3333
info threads
# * 1 Thread 1 "vexriscv.rv0" ...
# 2 Thread 2 "vexriscv.rv1" ... 0x00000000000000a4 in ?? ()
thread 2
p/x $mhartid # 0x1
For independent per-hart control, drop -rtos hwthread and target smp, then select a hart with
targets vexriscv.rv0 / targets vexriscv.rv1. Halting hart 1 leaves hart 0 running
(vexriscv.rv0 curstate → running, vexriscv.rv1 curstate → halted), and each hart steps and
resumes on its own.
DMI gateway vs Mem-AP (the RP2350 approach)
The closest production precedent is the Raspberry Pi RP2350: SWD pins, an ARM SW-DP, and a RISC-V Debug Module behind it (RP2350 datasheet §3.5, §3.8). Both designs are the same up to one layer. They differ only in the access port between the DP and the DM:
| RP2350 | This design (VexRiscv / VexiiRiscv) | |
|---|---|---|
| AP type | standard CoreSight APB Mem-AP, at 0x0a000 in the debug address space | custom 4-register designer AP — AP_IDR / DMI_ADDR / DMI_DATA / POSTED_READ (Phase 2C) |
Reaching DM register n | memory-mapped at n × 4: write the address to TAR, then access DRW (or BD0–BD3) | write n to DMI_ADDR, then read / write DMI_DATA |
| SWD packets per DM access | 1–4, depending on the access pattern | 1–2, depending on the access pattern |
| DP architecture | ADIv6 (SELECT holds the AP base address) | ADIv5 (APSEL field) |
| DP state at power-up | Dormant (needs the selection-alert wake-up sequence) | active |
| Standing in the RISC-V Debug Spec | custom DTM (Ch. 6) | custom DTM (Ch. 6) — the same |
Why this design keeps the gateway. What a Mem-AP would buy is abillity to speak Mem-AP lingo from ARM world which on the wire it is not cheap.
Speed
Both designs are an address register plus a data register — TAR + DRW on a Mem-AP, DMI_ADDR + DMI_DATA on the gateway However, the real difference is the Mem-AP's banked window. Host side tooling like OpenOCD's Mem-AP layer reads through BD0–BD3: TAR is set to a 16-byte-aligned address, and the four words of that window are then reached without touching TAR again. But BD0–BD3 sit in a different AP register bank from TAR, so moving to a new window costs a DP SELECT write, the TAR write, and a SELECT write back.
The gateway keeps all its registers in bank 0 and never writes SELECT.
This is explained below for SWD packets per DM access for both the two backends:
| Access | Gateway | Mem-AP (BD path) |
|---|---|---|
| Same register as the previous access | 1 | 1 |
| Another register in the same 4-register window | 2 | 1 |
| Register in a different window | 2 | 4 (SELECT + TAR + SELECT + BD) |
| End of a batch that read something | + 1 (RDBUFF) | + 1 (RDBUFF) |
Also RISC-V DM's registers cluster in those windows — dmcontrol / dmstatus, abstractcs /
command, and data0–data3 each share one — so for real operations:
| Operation | Gateway | Mem-AP |
|---|---|---|
Halt: write dmcontrol, poll dmstatus N times | 2 + 2 + (N − 1) | 4 + N |
Read a GPR, RV32: write command, poll abstractcs, read data0 | 7 | 10 |
Read a GPR, RV64: also read data1 | 9 | 11 |
| Poll the same register | 1 each | 1 each |
So on the wire. the gateway is a few packets ahead.
Area
In area the gateway is clearly smaller. Its AP is a 7-bit address register plus one DebugBus request path; the whole SWD DTM (PHY, DP, gateway, clock crossing) is 216 LUTs on an Artix-7. The SWD DTM breaks down as:
Part of DebugTransportModuleSwd | LUTs | FFs | Also needed with a Mem-AP? |
|---|---|---|---|
SwdPhy — wire protocol | 95 | 130 | yes, same DP |
SwdDp — DP registers | 25 | 50 | yes |
Response clock crossing (FlowCCByToggle) | 34 | 70 | yes — any AP needs a crossing to the DM (Hazard3 uses an async APB bridge) |
| Gateway AP + command clock crossing | 60 | 115 | no — this is the part a Mem-AP would replace |
| Total | 216 | 409 |
So the gateway costs at most 60 LUTs / 115 FFs, and part of that is the command-side clock crossing
a Mem-AP would need too. A Mem-AP in its place needs at least a 32-bit TAR, a CSW register
(access size, auto-increment, protection), the banked BD decode, an address incrementer and an
APB manager — roughly 100–150 LUTs by estimate (not synthesised).
Context - Mixed Architecture vs Pure RISC-V
RP2350 is a dual-architecture part: each core
slot holds an Arm Cortex-M33 and a Hazard3, selected at boot. Its debug complex is Arm CoreSight
around one SW-DP, with two AHB5 Mem-APs (debug address space 0x02000 / 0x04000) for the two
Cortex-M33s, the APB Mem-AP at 0x0a000 for the RISC-V DM, and RP-AP for chip-level control
(datasheet §3.5.2–3.5.3, Figure
6). The SW-DP and the Mem-AP infrastructure are there for the Arm cores anyway. And the Hazard3 DM
already exposes a byte-addressed APB port, as shown above. Connecting it through one more APB
Mem-AP therefore costs almost nothing extra, needs no new transport RTL, and makes the RISC-V
cores reachable through the same port, the same way, as the Arm ones.
This design starts from the opposite position: there is no Arm debug IP to reuse. The SW-DP is
written from scratch in SpinalHDL, and the SpinalHDL DM exposes a word-addressed DebugBus(7), not
APB. Building a Mem-AP would add a CoreSight-shaped register set and an APB manager only to reach a
DM that doesn't speak APB; the gateway reaches DebugBus directly. The one thing the Mem-AP shape
would buy is support in tools that only speak Mem-AP — and that is a host-side problem, solved here
by the vexriscv-gateway OpenOCD backend without touching the RTL.
| Mem-AP fits when | Gateway fits when | |
|---|---|---|
| Debug IP already on chip | an Arm CoreSight DAP is present (e.g. for Arm cores on the same die) | the DP is your own RTL |
| DM port | APB, byte-addressed (Hazard3) | word-addressed bus (DebugBus) |
| Host tooling | must work with tools that only drive Mem-APs | you control the host side (OpenOCD backend) |
| Area | the AP is already paid for | the smallest AP that does the job |
RISC-V-only chips: why the gateway is usually the better choice
Take the Arm cores away and RP2350's main reason for a Mem-AP disappears: there is no CoreSight debug IP on the die to reuse. For a SoC or MCU whose only processors are RISC-V, the gateway is usually the better fit:
- No CoreSight IP to reuse. The SW-DP has to be built anyway (here it is SpinalHDL). A Mem-AP
on top of it adds
TAR,CSW, the bankedBDdecode and an APB manager, with no functional gain: the DM is the only thing behind the AP. - The DM has a simple word-addressed port. SpinalHDL's
DebugModulespeaksDebugBus(7); the gateway connects to it directly, whereas a Mem-AP would first need an APB front end on the DM. - It is the smaller AP, and on the wire it is even or slightly ahead (see the tables above) — by tens of LUTs and a few packets per operation, so a tie-breaker rather than the main reason.
CPU-embedded debug plugin (EmbeddedRiscvJtag) over SWD
Both CPUs also have a plugin, EmbeddedRiscvJtag, that builds the Debug Module and its transport
inside a single-hart CPU. It is meant for small SoCs that instantiate the CPU directly. Examples are
VexiiRiscv's standalone Generate (--debug-jtag-*), VexRiscv's GenFullWithOfficialRiscvDebug and
Briey, and third-party SoCs such as aesc-silicon's nafarr / ElemRV. The plugin's SWD mode reuses
the same DebugTransportModuleSwd as the SMP cluster and the SoC-level debug fiber:
| PR | Option | Hardware check |
|---|---|---|
VexiiRiscv #188 — merged 2026-09-28 (19b41a7) | withSwd on the plugin, --debug-swd | Arty A7-100T, RV32, Black Magic Probe: load, software and single-step lane identical to the MCU-Link values |
VexRiscv #500 — merged 2026-09-27 (aefc0e0) | withSwd on the plugin | Arty A7-35T, LiteX "standard" core + official debug: MCU-Link (65 KB/s load, software + hardware breakpoints, ndmreset) and Black Magic Probe, identical results |
On both merges the existing JTAG configurations generate netlists identical to dev, and the SWD blocks are
byte-identical to the SMP cluster's. The host side (probe, OpenOCD fork, configs) is unchanged. The plugin
supports one hart; multi-hart designs use the SoC-level paths above.
ElemRV
elements-nafarr#72 adds SWD debug transport to ElemRV using EmbeddedRiscvJtag.
elements-zibal#62 passes it through the zibal platforms: a board sets debugTransport = DebugTransport.Swd in Hydrogen.Parameter (or Carbon / Nitrogen) and wires two pads, swclk and swdio, instead of four JTAG pads.
Black Magic Probe (OpenOCD alternative)
A Black Magic Probe (BMP) runs the GDB server on the probe: GDB
connects straight to its USB serial port, with no OpenOCD in between. Stock BMP firmware reads
this design's SW-DP but cannot use the DMI gateway AP. Support is proposed in
blackmagic #2322
(riscv_adi_dtm: support the SpinalHDL SWD "DMI gateway" AP; branch
disdi:feature/riscv-swd-dmi-gateway). No RTL change was needed.
Wiring (BMP v2.3, same Arty Pmod JB harness)
| BMP 10-pin | Signal | Arty |
|---|---|---|
| 1 | VTref (sets the probe's I/O level) | JB12 (3.3 V) |
| 2 | SWDIO | JB7 |
| 4 | SWCLK | JB3 |
| 3 / 5 / 9 | GND | JB11, plus the second ground JP2.3 → JB5 |
Hardware results
| Board / image | Result |
|---|---|
| Arty A7-35T, VexRiscv SMP (RV32IMA), golden image | rv32ima (exts 00141101). load demo.elf (6552 B), break main 0x4000069c, stepi → 0x400006b0, break help 0x40000644, bt #0 help / #1 0x400006cc main — identical to the OpenOCD + MCU-Link lane; compare-sections matches |
| Arty A7-100T, VexiiRiscv RV64 × 2 harts | DM v0.13, both harts found as two targets; attach, registers, single step, resume/halt per hart |
Use
# Probe firmware (BMP v2.x), from the blackmagic tree with #2322 :
meson setup build-fw --cross-file cross-file/bmp-v1-v2-riscv.ini && ninja -C build-fw
dfu-util -d 1d50:6018,:6017 -s 0x08002000:leave -D build-fw/blackmagic_bmp_v1_v2_firmware.bin
# GDB, straight to the probe (no OpenOCD):
riscv64-unknown-elf-gdb -ex 'set arch riscv:rv32' -ex 'set mem inaccessible-by-default off' \
-ex 'target extended-remote /dev/ttyACM0' -ex 'monitor swdp_scan' -ex 'attach 1' demo/demo.elf
(gdb) load
(gdb) break main
(gdb) continue
Testing
Phase 4: Verification and Testing
Host-visible sim paths that already pass are tracked under Phase 3
(raw-AP smoke, riscv examine, GDB attach/regs, demo break/continue). This chapter is
the broader regression matrix — much of it still open, especially on hardware.
Done in sim (via Phase 3)
- SWD-DTM →
DebugBus→ DM path at sbt (Phases 2A–2C) and LiteX Verilator SoC - OpenOCD master smoke: DPIDR +
dap apreg→dmstatus - OpenOCD
riscvexamine + halt/resume/register access over SWD (patched builds) - GDB attach / registers / memory R/W over SWD
- GDB
break main/continue/ backtrace on preloaded demo (no GDBload)
Still open
- Automated SWD stimuli via custom OpenOCD target on real CMSIS-DAP hardware (Arty)
- Test
dmactiveactivation/deactivation sequence (formal matrix) - Test abstract command error handling (all 7
cmderrcodes) - Test EBREAK behavior with
dcsr.ebreakm/ebreaks/ebreakucombinations - Test single-step across privilege mode transitions
- Test trigger module: each trigger type, chaining,
dmodesecurity - Test System Bus Access error handling (if implemented)
- Test authentication mechanism (if implemented)
- Decide
dmstatus.versionclaim (2= 0.13 vs3= 1.0) - Regression suite for all implemented features (CI-friendly)
Documentation
Phase 5: Documentation & Developer Interface
Done
- Phase 2A–2C architecture chapters in this book (
SwdPhy/SwdDp/SwdDmiGateway) - Operator how-to for sim three-terminal attach (JTAG and SWD) — Phase 3
- Operator how-to for Arty hardware (JTAG via BSCANE2 USER4; SWD via MCU-Link CMSIS-DAP on Pmod JB) — Phase 3
- Document placeholder
DPIDR,AP_IDR, and Phase 2CDMI_ADDR/DMI_DATA/POSTED_READmap (see Phase 2C) - Published OpenOCD Vexriscv fork: disdi/openocd
vexriscv-gateway(9786 + fixes + gateway; still unmerged upstream)
Still open
- Broader user-facing guide beyond this stack (probe-rs, multi-probe notes)
- Finalize published identification constants once IDs leave placeholder status
- Document which optional RISC-V debug-spec features are implemented vs stubbed
(
dmstatus.version, abstract commands surface, triggers, SBA)