Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Standard Debug Spec Milestone

Below are milestones planned in the project:

Status snapshot

AreaStatus
Phase 2A–2C (target-side SWD DTM)✅ Implemented and sim-verified (sbt + LiteX)
Phase 3 step 1 (LiteX + raw-AP smoke)✅ DPIDR + dmstatus over SWD
Phase 3 step 2 (riscv + GDB)✅ Vexriscv fork — disdi/openocd vexriscv-gateway (master + Gerrit 9786 + VexRiscv gateway)
Phase 3 step 2b (demo break / continue)✅ Preload path on the Vexriscv fork; no GDB load in sim
Phase 3 step 3 (JTAG on Arty)✅ 2026-08-20 — tunneled DTM via BSCANE2 USER4; on-board FT2232. Operator how-to: Phase 3
Phase 3 step 3 (SWD on Arty)✅ 2026-08-29 — MCU-Link CMSIS-DAP → Pmod JB; smoke + riscv examine + GDB load / break / stepi. Operator how-to: Phase 3
Phase 4–5Partial — sim and Arty attach done; full regression matrix and placeholder DPIDR / AP_IDR still open

Conformance note: the RISC-V Debug Specification defines only a JTAG DTM. SWD is a custom DTM (explicitly permitted). Claim "RISC-V Debug Specification, with custom DTM" — never an unqualified conformance claim.

Phase 1 adds missing features in DebugModule from the ratified RISC-V Debug Specification.

Phase 2 adds a target-side Arm Serial Wire Debug (SWD) front-end that still drives the same RISC-V DebugBus used by the JTAG DTM. The Arm specification is at ADIv5.0–ADIv5.2 (IHI0031).

Host (SWCLK / SWDIO)
        │
   ┌────▼────┐
   │ SwdPhy  │  Phase 2A — wire protocol (framing, turnaround, line reset)
   └────┬────┘
        │ SwdDpCmd / SwdDpRsp / SwdDpWrite
   ┌────▼────┐
   │  SwdDp  │  Phase 2B — DP registers + ACK policy (OK / WAIT / FAULT)
   └────┬────┘
        │ SwdApCmd / SwdApRsp
   ┌────▼──────────┐
   │ SwdDmiGateway │  Phase 2C — custom DMI AP + SWCLK↔debug CDC
   └────┬──────────┘
        │ DebugBus (same as JTAG DTM)
   ┌────▼────────┐
   │ DebugModule │  Phases 1B–1D
   └─────────────┘

How the SWD protocol works

What SWD is

Serial Wire Debug is a 2-wire, synchronous, packet-based host↔target link:

SignalRole
SWCLKClock (usually driven by the probe/host)
SWDIOBidirectional data (one bit per clock)

There is no separate reset pin. Recovery uses a line reset on SWDIO.

Logically, SWD talks to a Debug Port (DP). The DP either answers for itself (DP registers) or forwards accesses to an Access Port (AP) that reaches the real debug resource. In this project the AP is a custom RISC-V DMI gateway, which drives the existing DebugBus into the Debug Module.

One transaction on the wire

Every SWD access is request → ACK → (optional data).

Packet request (host → target, 8 bits)

After optional idle cycles (SWDIO low), the host sends:

Bit(s)NameMeaning
1StartAlways 1
1APnDP0 = DP register, 1 = AP register
1RnW0 = write, 1 = read
2A[2:3]Register address bits A[3:2] (LSB-first on the wire)
1ParityEven parity over APnDP, RnW, A[2], A[3]
1StopAlways 0
1ParkHost drives 1, then releases the line

All multi-bit fields are LSB first.

Turnaround (Trn)

When drive ownership changes, neither side should drive for a short period (default 1 SWCLK). That prevents bus fight on SWDIO.

ACK (target → host, 3 bits)

Encoding (value)NameWire order (LSB first)
0b001OK1, 0, 0
0b010WAIT0, 1, 0
0b100FAULT0, 0, 1

The numeric value and the bit order on the wire are easy to confuse: OK is value 001, so the first driven bit is 1.

Data phase (only if ACK = OK in this implementation)

DirectionWhenFormat
WriteAfter OK + second turnaround32-bit WDATA + even parity (host → target)
ReadAfter OK, no turnaround (target keeps the line)32-bit RDATA + even parity (target → host)

Errors and special sequences

ConditionTarget behaviour
Protocol error (bad request parity / Stop / Park)Do not drive ACK; stay silent until line reset
WDATA parity failACK already sent; DP sets WDATAERR and drops the write (not a line silence)
WAITBusy (e.g. previous AP access outstanding); host typically retries
FAULTSticky error set; host clears via ABORT
Line resetSWDIO HIGH for ≥ 50 SWCLK, then ≥ 2 idle (LOW)
IdleSWDIO low between frames, or back-to-back Start with zero idle

Clock and pins

Clock domain = SWCLK (probe-driven)

swdio.i  ← host/probe (or pad input)
swdio.o  → pad output value when target drives
swdio.oe → pad output enable
  • Target samples i and updates o/oe on rising SWCLK (OpenOCD bitbang model).

Minimal mental model (one frame)

  1. Host clocks an 8-bit header (Start…Park).
  2. Phy validates it; if bad → silence until line reset.
  3. If good → phy fires cmd and needs rsp.ack for the next few clocks.
  4. Phy drives ACK (value chosen by Phase 2B).
  5. If OK and read → phy streams rdata + parity from the DP.
  6. If OK and write → phy releases the line, samples wdata + parity, fires wr.
  7. If WAIT/FAULT → phy stops after ACK (no data phase here).
  8. Host may retry, clear sticky via ABORT or line-reset.

Debug Transport Module

Phase 1a: JTAG System Integration


Spec Feature Coverage Summary

Spec Feature (Chapter)Vexriscv Debug ImplementationImplementation Details
Ch 6: Debug Transport Module (DTM)
JTAG TAP with IDCODE/dtmcs/dmi✅ ImplementedDebugTransportModuleJtag.scala — standard IR codes, dtmcs with version/abits/idle/dmistat, dmi with op/address/data
DMI bus protocol✅ ImplementedDebugInterfaces.scala — DebugBus with DebugCmd/DebugRsp, DebugBusSlaveFactory
DMI busy/error handling✅ Implementeddmihardreset, dmireset, pending/overrun detection
Cross-clock-domain DMI✅ ImplementedccToggle for JTAG↔debug clock domains
JTAG tunnel support✅ ImplementedJtagTunnel.scala — tunneling through outer TAP

Debug Module

Phase 1B: Core DM Registers


Spec Feature (Chapter)Vexriscv Debug ImplementationImplementation Details
Ch 3: Debug Module (DM)
dmcontrol (0x10)✅ Implementeddmactive, ndmreset, haltreq, resumereq, ackhavereset, hartsello/hi. Missing: hasel, setresethaltreq, hartreset, keepalive
dmstatus (0x11)✅ Implementedversion, authenticated=1, all halted/running/unavail/nonexistent/resumeack/havereset flags, impebreak=1
hartinfo (0x12)✅ Implementeddataaddr=0, datasize=0, dataaccess=0, nscratch=0
abstractcs (0x16)✅ Implementeddatacount, progbufsize, busy, cmderr (all 7 error codes)
command (0x17) — Access Register✅ Implementedcmdtype=0 with full FSM: transfer, write, postexec, aarsize validation, GPR+FPU register access
command (0x17) — Access Memory❌ Not implementedReturns NOT_SUPPORTED
command (0x17) — Quick Access❌ Not implementedReturns NOT_SUPPORTED
abstractauto (0x18)✅ Implementedautoexecdata and autoexecProgbuf for burst access
progbuf0-N (0x20+)✅ ImplementedParameterized progbuf memory, multi-word execution with counter, redo support
data0-N (0x04+)✅ ImplementedMemory-backed, hart writes via fromHarts, host reads async
sbcs (0x38) — System Bus Access✅ Implemented (optional)sbversion=1, sbaccess, sbbusyerror, sbbusy, sbreadonaddr, sbautoincrement, sbreadondata, sberror, 32-bit only
sbaddress0 (0x39)✅ ImplementedRead/write with auto-increment
sbdata0 (0x3c)✅ ImplementedRead/write with bus triggers
sbaddress1-3 / sbdata1-3❌ Not implemented32-bit address/data only
sbcs 8/16/64/128-bit access❌ Not implementedOnly sbaccess32 supported
haltsum0 (0x40)✅ ImplementedPer-hart halted bits, up to 32 harts
haltsum1-3❌ Not implementedOnly haltsum0 exists
authdata (0x30)❌ Not implementedAlways authenticated
confstrptr0-3❌ Not implemented—
nextdm (0x1d)❌ Not implemented—
dmcs2 (0x32)❌ Not implemented—
Hart arrays (hawindowsel/hawindow)❌ Not implemented—
custom0-15❌ Not implemented—
Multi-hart support✅ ImplementedParameterized p.harts, per-hart buses, hartSel selection

CSR Register

Phase 1C: Debug CSRs (dcsr, dpc)


Spec Feature (Chapter)Vexriscv Debug ImplementationImplementation Details
Ch 4: Core Debug (hart-side CSRs)
Halt/Resume✅ DebugHartBus + CsrPluginCsrPlugin.scala (line 702+): running flag, DebugHartBus wiring, halt/resume handshake
Single-step✅ ImplementedCsrPlugin.scala (lines 807-845): dcsr.step with full FSM (IDLE→SINGLE→WAIT), timeout/redo handling
dcsr (0x7B0)✅ ImplementedCsrPlugin.scala (lines 792-851): prv, step, nmip, mprven, cause, stoptime, stopcount, stepie, ebreakm/s/u, xdebugver=4
dpc (0x7B1)✅ ImplementedCsrPlugin.scala (line 791): Reg(UInt(32 bits)), read/write via rw(CSR.DPC, dpc)
dscratch0 (0x7B2)❌ Not implemented—
dscratch1 (0x7B3)❌ Not implemented—
Debug mode entry✅ ImplementedCsrPlugin.scala (lines 1426-1443): saves PC→dpc, sets dcsr.cause (1=ebreak, 3=haltreq, 4=step), saves privilege→dcsr.prv, enters M-mode
Debug mode exit (resume)✅ ImplementedCsrPlugin.scala (lines 1488-1498): jumps to dpc, restores privilege from dcsr.prv, via DebugHartBus.resume
Halt cause reporting✅ Implementeddcsr.cause: 1 (ebreak), 2 (trigger), 3 (haltreq), 4 (step)
dcsr.ebreakm/s/u✅ ImplementedCsrPlugin.scala (lines 1372-1377): per-privilege ebreak→debug detection
dcsr.stoptime✅ ImplementedCsrPlugin.scala (line 866): stoptime output gated by debugMode
dcsr.stopcount✅ ImplementedCsrPlugin.scala (line 1176): mcycle increment gated by !debugMode || !stopcount
dcsr.stepie✅ ImplementedCsrPlugin.scala (line 1315): interrupts cleared when step && !stepie
Interrupt inhibition in debug✅ ImplementedCsrPlugin.scala (line 721): inhibateInterrupts() when debugMode
CSR access protection (0x7Bx)✅ ImplementedCsrPlugin.scala (line 1718): blocks non-debug access to 0x7B0-0x7BF
DebugHartBus wiring✅ ImplementedCsrPlugin.scala (lines 704-789): instruction injection, data CSR, all hartToDm/dmToHart signals
Reset control✅ dmcontrol.ndmreset → io.ndmresetBoth implementations provide ndmreset

Triggers

Phase 1D: Trigger Module


Spec Feature (Chapter)Vexriscv Debug ImplementationImplementation Details
Ch 5: Trigger Module (hart-side CSRs)
tselect (0x7A0)✅ ImplementedCsrPlugin.scala (lines 870-875): WARL index, parameterized debugTriggers (default 2)
tinfo (0x7A4)✅ ImplementedCsrPlugin.scala (line 878): reports type 2 (mcontrol) support
tdata1 (0x7A1)⚠️ PartialCsrPlugin.scala (lines 922-942): type=2, dmode, execute, m/s/u, action. Missing: timing, select, sizelo/hi, maskmax, chain, match, load, store, hit
tdata2 (0x7A2)✅ ImplementedCsrPlugin.scala (lines 944-953): 32-bit compare value, PC equality match
tdata3 (0x7A3)❌ Not implemented—
tcontrol (0x7A5)❌ Not implementedNo mte/mpte
Trigger type⚠️ mcontrol (type 2) onlyLegacy type 2, not type 6 (mcontrol6). Spec recommends type 6 for new implementations
Match modes⚠️ Equal onlyOnly match=0 (equality). No napot, >=, <, mask modes
Match targets⚠️ Execute address onlyOnly execute bit implemented. No load/store data/address match
Privilege filtering✅ Implementedm, s, u bits with privilegeHit logic (lines 930-934)
dmode security✅ Implementeddmode bit controls debug-only write access (line 925)
action field⚠️ PartialRegister exists but only action=1 (enter debug) used in match logic
Trigger chaining❌ Not implementedNo chain bit
dcsr.cause=2 on trigger✅ ImplementedCsrPlugin.scala (line 892): sets dcsr.cause := 2 on trigger hit
Trigger hit → debug entry✅ ImplementedCsrPlugin.scala (lines 881-897): decodeBreak halts pipeline, enters debug mode
mcontrol6 (type 6)❌ Not implementedOnly legacy type 2 exists
icount (type 3)❌ Not implemented—
itrigger/etrigger/tmexttrigger❌ Not implemented—
Hardware breakpoints⚠️ Spec-compliant but limitedType 2 mcontrol with execute address match=0 only, privilege filtering, dmode security
mcontext/scontext❌ Not implemented—
SoC Integration
DTM→DM→Hart wiring✅ ImplementedDebugModuleFiber.scala — multi-hart binding, clock-domain-safe pipelining
Tilelink SBA bridge✅ ImplementedmakeSysbusTilelink() in DebugModuleFiber

Phase 2A — SWD Protocol State Machine: Implementation & Verification

Code Repository :

Update submodules in pythondata-cpu-vexriscv_smp to below :

ArtifactPath
RTLEXT/SpinalHDL/lib/src/main/scala/spinal/lib/cpu/riscv/debug/DebugTransportModuleSwd.scala
TestbenchEXT/VexRiscv/src/test/scala/vexriscv/DebugSwdTest.scala
Runcd EXT/VexRiscv && sbt -batch "testOnly vexriscv.DebugSwdTest"

EXT = pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext. EXT/VexRiscv/build.sbt compiles EXT/SpinalHDL from source, so the RTL and testbench build in one sbt project with no LiteX involvement.


1. What is implemented — SwdPhy

SwdPhy is the ARM ADI SW-DP line layer (ADIv6.0 §B4): it speaks the 2-wire protocol and terminates in a decoded-transaction seam. It contains no DP registers (Phase 2B) and no DebugBus bridge (Phase 2C).

1.1 I/O

swdio.i  : in  Bool   -- SWDIO as driven by the probe
swdio.o  : out Bool   -- SWDIO value when the target drives
swdio.oe : out Bool   -- target output enable
dp.cmd   : master Flow(SwdDpCmd)    -- decoded request        (2A -> 2B)
dp.rsp   : slave  Flow(SwdDpRsp)    -- ACK + read data        (2B -> 2A)
dp.wr    : master Flow(SwdDpWrite)  -- write commit           (2A -> 2B)
  • Clock = SWCLK. The component's implicit clock domain is the probe-driven SWCLK. The target samples swdio.i and updates swdio.o/swdio.oe on the rising edge, matching the OpenOCD bitbang model (host sets data while SWCLK is low, samples target data while low).
  • No inout. The tristate split (i/o/oe) is required because the cluster is a Verilog black box in LiteX.

1.2 The 2A↔2B seam (three flows)

A wire-protocol fact shapes the seam: a write's ACK is sent before the 33-bit WDATA phase, so write data cannot ride in the request.

FlowFiredPayload
SwdDpCmdone-cycle pulse on the packet-request park edge, iff parity/stop/park all passapNdp, rnw, addr = A[3:2]
SwdDpRspmust be presented by the DP within the turnaround cycle (an always-ready DP may simply hold it valid)ack (OK=001/WAIT=010/FAULT=100), rdata
SwdDpWriteone-cycle pulse after the 33rd WDATA bitdata, parityOk (false ⇒ WDATAERR material for 2B)

The response is latched into a hold register on first sight (rsp.valid may be combinational off cmd or held continuously); ACK and RDATA are driven from the latched copy for the rest of the frame.

Read path — READ_DATA → RELEASE → IDLE (Fig B4-2)

  • 32 data bits LSB-first from rspHold.rdata, then even parity (xorR).
  • No turnaround between ACK and RDATA (target keeps oe=1).
  • RELEASE: oDrive := False, then IDLE (trailing Trn / release).

Write path — WR_TRN → WRITE_DATA → IDLE (Fig B4-1)

  • WR_TRN: two cycles with oe=0 (cnt 0 then 1) = second turnaround.
  • WRITE_DATA: shift in 32 bits + sample parity; fire dp.wr with data and parityOk.
  • WDATA parity fail → parityOk = false (WDATAERR material for 2B), not protocol error.
  • Return direct WRITE_DATA → IDLE (no RELEASE): host already owns the line. Diagram §1.4.2 also uses WRITE_DATA --> IDLE.

ERROR and line reset

Diagram §1.4.2RTL
ERROR --> ERROR (ignore traffic)ERROR: oDrive := False only; no header parse
ERROR --> IDLE on line resetlineReset.hit → RESET_WAIT → first low → IDLE

Line reset is orthogonal (overrides any state): ≥50 consecutive highs on swdio.i while !oDrive; counter frozen/cleared while target drives so ACK/RDATA cannot fake a reset.

RESET_WAIT is an RTL-only gate so the target does not accept Start until the line has gone idle after the reset burst. Diagram §1.4.2 draws a direct ERROR → IDLE; recovery contract is the same.


2. How the testbench works — DebugSwdTest

The bench is the probe

SWCLK is the DUT clock, and the bench owns it: no forkStimulus — every SWCLK cycle is one call to step(bit):

fallingEdge(); set swdio.i = bit;      // host updates while SWCLK low
sample (swdio.o, swdio.oe);            // host samples while SWCLK low
risingEdge();                          // target samples/updates

This reproduces OpenOCD's bitbang_swd_exchange exactly: a host-driven bit is sampled by the target at the rising edge ending its cycle; a target-driven bit read in cycle k is the value the target registered at edge k−1. All multi-bit values are sent/collected LSB first. step returns (o, oe), so every helper can assert drive/release behavior per cycle.

Phase 2B — SW-DP Register File: Implementation & Verification

Code Repository :

Update submodules in pythondata-cpu-vexriscv_smp to below :

ArtifactPath
RTL (SwdDp, SwdPhyDp, AP seam bundles)EXT/SpinalHDL/lib/src/main/scala/spinal/lib/cpu/riscv/debug/DebugTransportModuleSwd.scala (same file as Phase 2A)
Shared probe driverEXT/VexRiscv/src/test/scala/vexriscv/SwdSimDriver.scala
TestbenchEXT/VexRiscv/src/test/scala/vexriscv/DebugSwdDpTest.scala
Runcd EXT/VexRiscv && sbt -batch "testOnly vexriscv.DebugSwdDpTest" (both phases: "testOnly vexriscv.DebugSwdTest vexriscv.DebugSwdDpTest")

EXT = pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext. Still zero LiteX involvement.


1. What is implemented

1.1 SwdDp — the ADI SW-DP register file

SwdDp sits behind the Phase 2A seam (SwdDpCmd/SwdDpRsp/SwdDpWrite, consumed as slave flows) and exposes a new AP seam toward Phase 2C:

ap.cmd : master Flow(SwdApCmd)   -- rnw, addr = A[3:2], apSel = SELECT[31:24], wdata
ap.rsp : slave  Flow(SwdApRsp)   -- error, data (completion; test-controlled latency)

1.2 DP register map (SWD A[3:2] encoding)

A[3:2]ReadWrite
00DPIDR (parameter; default 0x0BA11AAB placeholder)ABORT — DAPABORT, STKCMPCLR, STKERRCLR, WDERRCLR, ORUNERRCLR
01CTRL/STAT when DPBANKSEL==0; other banks read as zeroCTRL/STAT when DPBANKSEL==0; other banks write-ignored
10RESEND (= last posted result)SELECT — APSEL[31:24], APBANKSEL[7:4], DPBANKSEL[3:0]
11RDBUFF (= last posted result)TARGETSEL (SWD v2) — accepted, ignored

These two diagrams put it back on the wire, making explicit where each step of a DP access is accomplished: the frame shape is Phase 2A (SwdPhy), the register semantics are Phase 2B (SwdDp). Neither phase performs a DP access on its own.

A DP read — DPIDR, A[3:2] = 00

Wire layout — one cell per SWCLK bit period, LSB first within every multi-bit field:

Bits 0–7 host-driven · bits 9–44 target-driven · bits 8 and 45 are turnarounds where neither side drives and SWDIO floats to its mandatory pull-up. The closing turnaround is one bit period on the wire but two SWCLK cycles in RTL (RELEASE), because SwdPhy registers its outputs — see the step-ownership table below.

Phase ownership:

A DP write — SELECT, A[3:2] = 10

Wire layout — note the second turnaround at bit 12, absent from the read:

Bits 0–7 host-driven · bits 9–11 target-driven · bits 13–45 host-driven again · bits 8 and 12 are turnarounds. This layout is the reason the 2A↔2B seam has three flows: the target must commit to the ACK at bits 9–11, thirty-two bit periods before the data it is acknowledging exists on the wire.

Phase ownership:


2. Testbench changes

2.1 Shared driver extraction

The probe-side bit-bang logic (step, header, readAck, transactRead, transactWrite, lineReset, with all embedded protocol assertions) moved from the 1a suite into SwdSimDriver.scala (SwdHostDriver, SwdAckSim). DebugSwdTest delegates to it with its 9 test bodies unchanged — proven by the suites running together (18/18). Step 2 will reuse the same driver against the full transport + DebugModule.

2.2 The stub moves back one layer

Exactly as the step-1b plan prescribes: the always-ready DP stub is gone — the real SwdDp answers within the turnaround by construction (combinational response). The test harness now stubs the AP side:

  • an observer thread records every ap.cmd fire into a queue ((rnw, addr, wdata));
  • apComplete(data, error) delivers a completion on ap.rsp for exactly one SWCLK cycle — completions are test-controlled, which is what makes posted-read ordering and WAIT-while-busy directly testable (the AP simply doesn't respond until the test says so).

Sugar wrappers keep tests readable: dpRead/dpWrite/apRead/apWrite map to driver transactions with APnDP set accordingly.


4. Verification

cd ~/fpga/pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext/VexRiscv
sbt -batch "testOnly vexriscv.DebugSwdTest vexriscv.DebugSwdDpTest"
# expect: Tests: succeeded 18, failed 0, canceled 0, ignored 0, pending 0

Phase 2C — DMI Gateway, Clock Crossing & DebugModule Integration

Code Repository :

OR

Update submodules in pythondata-cpu-vexriscv_smp to below :

ArtifactPath
RTL (SwdDmiGateway, DebugTransportModuleSwd)EXT/SpinalHDL/lib/src/main/scala/spinal/lib/cpu/riscv/debug/DebugTransportModuleSwd.scala (same file as 2A/2B)
Fiber hook (withSwdTransport())EXT/SpinalHDL/lib/src/main/scala/spinal/lib/cpu/riscv/debug/DebugModuleFiber.scala
Testbench (incl. SwdDmTestTop)EXT/VexRiscv/src/test/scala/vexriscv/DebugSwdDmTest.scala
Runcd EXT/VexRiscv && sbt -batch "testOnly vexriscv.DebugSwdDmTest" (all: "testOnly vexriscv.DebugSwdTest vexriscv.DebugSwdDpTest vexriscv.DebugSwdDmTest")

EXT = pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext. Still zero LiteX involvement.


1.Implementation of DMI for SWD

1.1 SwdDmiGateway — the RISC-V DMI gateway AP

Sits behind the 2B AP seam (SwdApCmd/SwdApRsp) and drives a DebugBus master — the same DebugCmd/DebugRsp interface the JTAG DTM produces, which is the whole point of the transport-agnostic design.

The DMI needs a 7-bit address and 32 bits of data. SWD offers, per transaction, one bit of DP/AP selection and two bits of register address. Four AP registers, total.

AP A[3:2]RegisterAccessBehavior
00AP_IDRROidentification constant (default 0x74726976 "triv") — completes locally in 1 SWCLK cycle
01DMI_ADDRRWlatches the 7-bit DMI word address; readable back — local
10DMI_DATARWperforms the DebugBus transaction at DMI_ADDR (RnW from the SWD packet)
11POSTED_READROlast completed DMI read result — local

Writes to RO registers complete OK with no effect. SELECT.APSEL is not decoded (single-AP design). AP reads launch at the request (posted, per 2B); AP writes launch at the WDATA commit carrying the data.

1.2 DebugTransportModuleSwd — the full transport

SWD pins → SwdPhy (2A) → SwdDp (2B) → SwdDmiGateway (2C) → DebugBus. The SWD side elaborates under ClockDomain(swclk, resetKind = BOOT); the DebugBus side under the provided debugCd. This is the SWD counterpart of DebugTransportModuleJtagTap.

DebugModuleFiber.withSwdTransport(dpidr, apIdr) instantiates it alongside withJtagTap() — same one-transport-per-build rule (direct io.ctrl connection, no arbitration).

1.3 End-to-end dmstatus read (the step-2 exit), graphically


2. Testbench — DebugSwdDmTest

  • DUT = SwdDmTestTop: the full transport + a real DebugModule (version 2, 1 hart, progBufSize=2, datacount=1) with a stubbed DebugHartBus — running, never halted, hartToDm/resume.rsp idle.
  • Two genuinely asynchronous clocks: the shared SwdHostDriver bit-bangs SWCLK (bench-owned, via a ClockDomain handle over the pin) while dut.clockDomain free-runs via forkStimulus — DMI completion latency is variable by construction.
  • Host-style retry helpers: apWriteRetry/apReadRetry/rdbuff retry on WAIT with bounded attempts — the same loop a real OpenOCD target would run. dmiRead/dmiWrite compose them into DMI operations.
  • Failure forensics built in: on any unexpected ACK the harness reads CTRL/STAT and includes the sticky flags in the assertion message (ackCheck) — this is what cracked the phantom-completion bug (§4.2). The sim(name, seed = …) hook pins a failing seed for deterministic reproduction.

3. Verification

cd ~/fpga/pythondata-cpu-vexriscv-smp/pythondata_cpu_vexriscv_smp/verilog/ext/VexRiscv
sbt -batch "testOnly vexriscv.DebugSwdDmTest"
# expect: Tests: succeeded 7, failed 0, canceled 0, ignored 0, pending 0
sbt -batch "testOnly vexriscv.DebugSwdTest vexriscv.DebugSwdDpTest vexriscv.DebugSwdDmTest"
# expect: Tests: succeeded 25, failed 0, canceled 0, ignored 0, pending 0

4. Summary

Phases 2A–2C complete the target-side RTL of the SWD plan; the DebugBus now speaks SWD end-to-end in simulation.

System Integration — SWD

Goal is host-visible integration only (LiteX SoC, OpenOCD, GDB). Target-side SWD RTL (SwdPhy / SwdDp / SwdDmiGateway) is Phase 2A–2C and is not re-defined here.


Code repositories

PieceWhereStatus
SWD DTM (DebugTransportModuleSwd) + DebugModuleFiber.withSwdTransport()SpinalHDL #1956merged 2026-09-08
spinal.lib.com.swd split (Swd / SwdPhy / SwdDp)SpinalHDL #1966merged 2026-09-19
VexRiscv SMP cluster --swdVexRiscv #483, #499merged 2026-09-08 / 09-19
VexiiRiscv LiteX SoC --with-swd + MicroSoc --swdVexiiRiscv #184merged 2026-09-23
VexiiRiscv CPU-embedded debug plugin (EmbeddedRiscvJtag) over SWD, --debug-swdVexiiRiscv #188merged 2026-09-28 (19b41a7)
VexRiscv CPU-embedded debug plugin (EmbeddedRiscvJtag) over SWDVexRiscv #500merged 2026-09-27 (aefc0e0)
ElemRV / nafarr: DebugTransport (Jtag default / Swd) on the VexiiRiscv realtime and performance presetselements-nafarr #72merged 2026-09-29 (54406843). VexiiRiscv pin 19b41a7 is #74 (b7a1257)
ElemRV / zibal: debugTransport on the Hydrogen / Carbon / Nitrogen platforms → io_plat.jtag or io_plat.swdelements-zibal #62open, submitted 2026-09-29 (d78489b on main 8f5e3f5, which already pins nafarr 54406843)
JTAG on Xilinx USER chains (add_cpu_jtag_debug, --with-cpu-jtag-debug)LiteX #2572 + linux-on-litex-vexriscv #459merged 2026-09-10
LiteX: swdremote sim module, OpenOCD configs, --with-swd-debug for vexriscv_smphttps://github.com/disdi/litex/tree/swdbranch, not yet proposed upstream
linux-on-litex-vexriscv: SWD pads on Arty Pmod JBhttps://github.com/disdi/linux-on-litex-vexriscv/tree/swd-artybranch, waits for the LiteX part
OpenOCD (Vexriscv fork)https://github.com/disdi/openocd/tree/vexriscv-gatewaybranch; Gerrit 9786 + gateway backend
Black Magic Debug: DMI gateway AP support (no OpenOCD)blackmagic #2322 — disdi:feature/riscv-swd-dmi-gatewayopen, submitted 2026-09-25 — see Black Magic Probe

OpenOCD (host ONLY) — :

LaneBuildRole
raw-AP smokeOpenOCD master (stock OK for SWD remote_bitbang)DPIDR + dap apreg → dmstatus; no GDB
Vexriscv fork (riscv + GDB)disdi/openocd vexriscv-gateway — OpenOCD master + Gerrit 9786 + designer-AP / VexRiscv DTM backendexamine + halt/resume/regs + GDB :3333

Stock master is enough for smoke. Full riscv attach needs the published vexriscv-gateway branch: 9786 (DTM + Mem-AP DMI backend), two fixes to 9786 itself, and a second designer-AP / VexRiscv gateway backend for the Phase 2C DMI_ADDR / DMI_DATA map. No RTL change.

Upstream’s stock riscv target rejects -dap at argument parsing, so Tcl-only dap apreg helpers cannot drive a GDB session — that is why the 9786 + gateway path exists.


Simulation based workflow using verilator

Side-by-side

TerminalJTAG — full three-terminal ✅SWD — full three-terminal ✅
1 — simlitex_sim … --with-privileged-debug --jtag-tap --with-jtagremote → TCP 44853 (jtagremote)litex_sim … --with-privileged-debug --with-swd-debug --with-swdremote → TCP 44854 (swdremote); add --ram-init=demo.bin for demo debug
2 — OpenOCDStock riscv target + fabric TAP — no vendor BSCAN in sim; examines hart; GDB :3333transport select swd + DAP + Vexriscv fork riscv; examines hart; GDB :3333
3 — GDBtarget extended-remote localhost:3333 → halt / regs / loadattach + regs ✅; demo break main / continue / bt ✅ via preload (no GDB load)
CapabilityJTAGSWD
Verilator SoC + official DM✅✅ (_Swd cluster)
Wire transport in sim✅ JTAG TAP + tunnel✅ SW-DP (DebugTransportModuleSwd)
OpenOCD sees transport✅ TAP 0x10003fff✅ SWD DPIDR 0x0ba11aab
Read dmstatus✅ via riscv / DMI✅ via vexriscv_dmi_read 0x11 / smoke
Examined RISC-V core✅✅ XLEN=32, misa=0x40141101
GDB halt / resume / info registers✅✅
Break / continue / backtrace✅ (load OK on JTAG)✅ via --ram-init=demo.bin + symbols; GDB load impractical in sim
OpenOCD binarystock masterstock master for raw-AP smoke; Vexriscv fork (disdi/openocd vexriscv-gateway) for riscv / GDB

JTAG — end-to-end workflow

Official stack only (--with-privileged-debug + full JTAG TAP in sim).

Prerequisites

  • litex_sim (LiteX venv)
  • OpenOCD master with standard RISC-V target
  • riscv64-unknown-elf-gdb

Configs

FileRole
openocd_jtag_remote.cfgremote_bitbang → localhost:44853
riscv_jtag_tunneled.tclTAP irlen 6, ID 0x10003fff, riscv use_bscan_tunnel 6 1

Terminal 1 — sim (keep running)

litex_sim \
  --integrated-main-ram-size=0x10000 \
  --cpu-type=vexriscv_smp \
  --cpu-variant=linux \
  --cpu-count=1 \
  --with-privileged-debug \
  --jtag-tap \
  --with-jtagremote \
  --non-interactive
FlagRole
--integrated-main-ram-size=0x1000064 KiB main RAM for sim (demo load region)
--cpu-type=vexriscv_smp / --cpu-variant=linux / --cpu-count=1SMP Linux-capable cluster, 1 hart
--with-privileged-debugOfficial DebugModule + DTM (_Pd netlist token)
--jtag-tapFull JTAG TAP on cluster (_JtagT); needed so sim has TCK/TMS/TDI/TDO pads
--with-jtagremoteLiteX sim module jtagremote — OpenOCD remote_bitbang on TCP 44853
--non-interactiveKeep sim running (no local control menu); target for OpenOCD/GDB

Wait for Found port 44853 and BIOS prompt litex>. First run may take several minutes (cluster regen + Verilator compile).

Terminal 2 — OpenOCD (after Terminal 1 is up)

openocd -f openocd_jtag_remote.cfg -f riscv_jtag_tunneled.tcl

Success indicators:

Info : JTAG tap: riscv.cpu tap/device found: 0x10003fff
Info : Examined RISC-V core; found 1 harts
Ready for Remote Connections
Info : Listening on port 3333 for gdb connections

Terminal 3 — GDB (after OpenOCD is ready)

riscv64-unknown-elf-gdb demo/demo.elf
set remotetimeout 120
set pagination off
set arch riscv:rv32
target extended-remote localhost:3333
monitor reset halt
x/8i $pc
info registers

One-liner:

riscv64-unknown-elf-gdb -ex "set remotetimeout 120" \
  -ex "target extended-remote localhost:3333" \
  demo/demo.elf

Load demo.elf (linked at 0x40000000) only after halt — this litex_sim invocation does not pass --ram-init=demo.bin, so the image is not preloaded:

monitor reset halt
load demo/demo.elf
break main
continue

If sim is slow and keep_alive() warnings (slow bitbang) are seen, prefer target extended-remote and set remotetimeout 120.


SWD — OpenOCD

LaneOpenOCD buildConfigsGives you
raw APmasteropenocd_swd_remote.cfg + vexriscv_swd.cfgDPIDR + dap apreg → dmstatus; no GDB
riscv Vexriscv forkdisdi/openocd vexriscv-gateway (master + 9786 + gateway)+ vexriscv_swd_riscv_master.cfgexamine + halt/resume/regs + GDB :3333

Prerequisites (SWD-specific)

Configs

FileRole
openocd_swd_remote.cfgremote_bitbang → localhost:44854, transport select swd
vexriscv_swd.cfgSW-DP + DAP + vexriscv_dmi_read/write + vexriscv_swd_smoke — no riscv target
vexriscv_swd_riscv_master.cfgVexriscv fork: dtm create -type vexriscv-gateway + riscv + gdb-attach halt

SWD Reading dmstatus, all the way down

Terminal 1 — sim (keep running)

GoalExtra flag
Attach / regs / raw-AP smoke(none) — wait for Found port 44854 + BIOS litex>
Debug the demo app (break main / continue / bt)--ram-init=demo.bin — wait for serialboot timeout → Executing booted program at 0x40000000 → litex-demo-app>
litex_sim \
  --integrated-main-ram-size=0x10000 \
  --cpu-type=vexriscv_smp \
  --cpu-variant=linux \
  --cpu-count=1 \
  --with-privileged-debug \
  --with-swd-debug \
  --with-swdremote \
  --non-interactive
  --ram-init=demo.bin
FlagRole
--with-privileged-debugOfficial DebugModule (required; SWD is official-stack only)
--with-swd-debugCluster SWD transport + _Swd netlist token
--with-swdremoteLiteX sim module swdremote — OpenOCD SWD bitbang on TCP 44854
--ram-init=demo.binPreload demo into main_ram @ 0x40000000 (demo-debug only)

Terminal 2 — raw-AP smoke (stock master; no GDB)

One-shot:

openocd -s tcl \
  -f openocd_swd_remote.cfg \
  -f vexriscv_swd.cfg \
  -c init -c vexriscv_swd_smoke -c shutdown

Verified success:

Info : SWD DPIDR 0x0ba11aab
AP_IDR   = 0x74726976
dmstatus = 0x004c0c82 (version=2 authenticated=1 allrunning=1 allhalted=0)
PASS: SWD -> SW-DP -> DMI gateway -> DebugModule

Interactive Tcl helpers (same two configs, stay open):

vexriscv_swd_smoke
vexriscv_dmi_read 0x11          ;# dmstatus

Terminal 2 — riscv target + GDB server (Vexriscv fork: disdi/openocd vexriscv-gateway)

# Use the openocd binary built from:
#   https://github.com/disdi/openocd/tree/vexriscv-gateway
openocd -s tcl \
  -f openocd_swd_remote.cfg \
  -f vexriscv_swd.cfg \
  -f vexriscv_swd_riscv_master.cfg

Verified examine (real hart, not stub):

Info : SWD DPIDR 0x0ba11aab
Info : [vexriscv.rv] datacount=1 progbufsize=2
Info : [vexriscv.rv] Examined RISC-V core
Info : [vexriscv.rv]  XLEN=32, misa=0x40141101
vexriscv.rv halted due to debug-request.

misa=0x40141101 = RV32 I+M+A+S+U. Check halt/resume by curstate, not only by log lines: resume may print halted due to single-step. while stepping off a breakpoint — that is not a failure.

Terminal 3 — GDB (after the Vexriscv fork is listening on :3333)

Start GDB with the ELF for symbols (attach-only or demo-debug):

riscv64-unknown-elf-gdb demo/demo.elf

Attach and inspect

set remotetimeout 300
set pagination off
set arch riscv:rv32
target extended-remote localhost:3333
info registers
x/6i $pc

vexriscv_swd_riscv_master.cfg sets -event gdb-attach halt, so GDB attaches to an already-halted target — no monitor halt is required.

Debug the demo app (break / continue / bt)

If litex_sim is passed with --ram-init=demo.bin :

set remotetimeout 300
set pagination off
set arch riscv:rv32
target extended-remote localhost:3333
# Image is already in main_ram via --ram-init=demo.bin.

x/8xw 0x40000000        # confirm preload (e.g. 0x0b00006f 0x00000013 ...)
set $pc = 0x40000000    # re-enter demo at _start so main is hit cleanly
break main
continue
bt

Expected: stop at main (typically around 0x4000069c); bt shows #0 main ().


Hardware based workflow using Arty

Reference manual: https://digilent.com/reference/programmable-logic/arty-a7/reference-manual

Devicexc7a35ticsg324-1L (a7-35 variant)
System clockclk100, 100 MHz
USBOn-board FTDI FT2232HQ — USB-JTAG (programming) + USB-UART (console) on one cable
ProgrammerOpenOCD via openocd_xc7_ft2232.cfg + bscan_spi_xc7a35t.bit
PMOD connectorsJA/JB/JC/JD → pmoda/pmodb/pmodc/pmodd
SWD probeExternal CMSIS-DAP (MCU-Link) on Pmod JB — not the on-board FTDI

JTAG on Arty — end-to-end workflow

Official stack only (--with-privileged-debug).

--with-privileged-debug alone is not enough on hardware. Without --jtag-tap the DTM is tunneled, and its debugPort_* signals need a vendor boundary-scan primitive (BSCANE2 on a Xilinx USER chain). That binding is now upstream as an explicit second flag, --with-cpu-jtag-debug — LiteX #2572 (LiteXSoC.add_cpu_jtag_debug(), default USER4, IR 0x23) and linux-on-litex-vexriscv #459. USER1 stays free for jtagbone. It replaces the original #458, which was closed unmerged; the hardware results below were taken with #458's BSCANE2 instance, the same USER4 / IR 0x23 binding.

Prerequisites

  • Arty
  • OpenOCD master with standard RISC-V target
  • riscv64-unknown-elf-gdb

Configs

FileRole
openocd_arty_bscan.cfgFits on-board FT2232 into the Xilinx TAP
openocd_arty_official.cfgJTAG on Arty hardware. Composes openocd_arty_bscan.cfg (adapter + Xilinx TAP) with riscv_jtag_tunneled.tcl (the riscv target + use_bscan_tunnel).

Terminal 1 — Build and Flash on Arty

# Stock LiteX + linux-on-litex-vexriscv master (litex#2572 / linux-on-litex-vexriscv#459)
./make.py --board=arty --cpu-count=1 --with-privileged-debug --with-cpu-jtag-debug --build --load

Expected in OpenOCD: Examined RISC-V core; found 1 harts and XLEN=32, misa=0x40141101.

Terminal 2 — OpenOCD (after Terminal 1 is up)

openocd -f litex/litex/tools/debug/openocd_arty_official.cfg

Terminal 3 — GDB (after OpenOCD is ready)

riscv64-unknown-elf-gdb -ex "set arch riscv:rv32"   -ex "target extended-remote localhost:3333"   linux-on-litex-vexriscv/build/arty/software/bios/bios.elf

SWD on Arty — end-to-end workflow

Official stack only (--with-swd-debug; that flag implies --with-privileged-debug). One transport per bitstream — do not combine with --jtag-tap.

--with-swd-debug alone is not enough on hardware. Unlike JTAG (which reuses the on-board FT2232 as a tunneled TAP), SWD needs an external CMSIS-DAP probe and two user I/O pins. Pinout lives in soc_linux.py (_swd_pmod_io / add_cpu_swd_debug) on https://github.com/disdi/linux-on-litex-vexriscv/tree/swd-arty.

Prerequisites

  • Arty
  • CMSIS-DAP probe (MCU-Link is the verified one) + jumper wires
  • OpenOCD master for the raw-AP smoke; the Vexriscv fork (disdi/openocd vexriscv-gateway) for GDB support.
  • riscv64-unknown-elf-gdb

Hardware connection

Two USB cables to the host. The FTDI does not carry SWD.

Host PC
 ├─ USB ── FT2232 (Arty J10) ── bitstream load + UART console
 └─ USB ── MCU-Link (CMSIS-DAP)
              SWCLK ──► JB3 ──► cluster swd_clk ──► SwdPhy
              SWDIO ◄─► JB7 ── IOBUF (swdio_i / o / oe)
                                 └── SwdPhy → SwdDp → DMI gateway → DebugModule

--with-swd-debug brings SWCLK / SWDIO out on Pmod JB (high-speed header: no 200 Ω series resistors, which matters for bidirectional SWDIO turnaround). Do not use JA or JD.

SignalLiteX pinPmod JBFPGA ballNotes
SWCLKpmodb:2JB3D15Probe-driven, gated clock
SWDIOpmodb:4JB7J17Bidirectional; FPGA PULLUP TRUE (ADI)
GND—JB11—Common ground (required)
GND (2nd)—JB5—Required — second ground return, see below
3.3 V (VTref)—JB12—Required for the MCU-Link — see below

Looking into the 12-pin Pmod:

JB1        JB2  JB3=SWCLK  JB4   JB5=GND(2nd)  JB6=3V3
JB7=SWDIO  JB8  JB9        JB10  JB11=GND      JB12=3V3 (VTref)

MCU-Link 10-pin Cortex debug header → Arty:

MCU-Link pin 4 (SWCLK)    → Arty JB3
MCU-Link pin 2 (SWDIO)    → Arty JB7
MCU-Link pin 3 (GND)      → Arty JB11
MCU-Link pin 5 (GND)      → Arty JB5    (second ground, required)
MCU-Link pin 1 (VTref)    → Arty JB12   (required on the MCU-Link)

Configs

FileRole
openocd_arty_swd.cfgCMSIS-DAP adapter. Hardware sibling of openocd_swd_remote.cfg — same DAP/DMI helpers, remote_bitbang swapped for cmsis-dap + usb_bulk.
vexriscv_swd.cfgSW-DP + DAP + vexriscv_dmi_read/write + vexriscv_swd_smoke — reused from sim, unchanged
vexriscv_swd_riscv_master.cfgVexriscv fork: dtm create -type vexriscv-gateway + riscv + gdb-attach halt — reused from sim, unchanged

Terminal 1 — Build and Flash on Arty

#   Support for SWD added to https://github.com/disdi/linux-on-litex-vexriscv/tree/swd-arty
./make.py --board=arty --cpu-count=1 --with-swd-debug --build --load

The first SWD build regenerates the cluster via sbt (no _Swd netlist ships prebuilt). Expected name: VexRiscvLitexSmpCluster_Cc1_…_Ood_Pd_Hb1_Swd.

Terminal 2 — raw-AP smoke (stock master; no GDB)

openocd -s tcl \
  -f litex/litex/tools/debug/openocd_arty_swd.cfg \
  -f litex/litex/tools/debug/vexriscv_swd.cfg \
  -c init -c vexriscv_swd_smoke -c shutdown

Success indicators:

Info : SWD DPIDR 0x0ba11aab
AP_IDR   = 0x74726976
dmstatus = 0x004c0c82 (version=2 authenticated=1 allrunning=1 allhalted=0)
PASS: SWD -> SW-DP -> DMI gateway -> DebugModule

dmstatus is bit-identical to the sim value.

Terminal 2 — riscv target + GDB server (Vexriscv fork: disdi/openocd vexriscv-gateway)

# Use the openocd binary built from:
#   https://github.com/disdi/openocd/tree/vexriscv-gateway
openocd -s tcl \
  -f litex/litex/tools/debug/openocd_arty_swd.cfg \
  -f litex/litex/tools/debug/vexriscv_swd.cfg \
  -f litex/litex/tools/debug/vexriscv_swd_riscv_master.cfg

Success indicators:

Info : SWD DPIDR 0x0ba11aab
Info : [vexriscv.rv] Examined RISC-V core
Info : [vexriscv.rv]  XLEN=32, misa=0x40141101

misa=0x40141101 = RV32 I+M+A+S+U.

Terminal 3 — GDB (after OpenOCD is ready)

riscv64-unknown-elf-gdb -ex "set arch riscv:rv32"   -ex "target extended-remote localhost:3333"   linux-on-litex-vexriscv/build/arty/software/bios/bios.elf

Hardware is not limited the way sim is: GDB load and stepi both work (CMSIS-DAP at 1 MHz vs swdremote pacing). Debug demo.elf (linked at 0x40000000):

monitor halt
load demo/demo.elf
set $pc = 0x40000000
break main
continue
bt

Expected: stop at main (typically around 0x4000069c); bt shows #0 main ().


VexiiRiscv over SWD

The same SWD transport and DebugModule now also serve VexiiRiscv VexiiRiscv#184. No new transport RTL was needed. VexiiRiscv builds its debug logic from SpinalHDL's DebugModuleSocFiber, and the SWD DTM is added in that fiber's body with dm.withSwdTransport(). SwdPhy / SwdDp / SwdPhyDp in the generated netlist are byte-identical to the VexRiscv SMP cluster's, so everything on the host side (probe, OpenOCD fork, configs) is reused unchanged. The SWD transport is the same SpinalHDL code in both CPUs, and the generated Verilog confirms it.

SoCOptionTop-level ports
LiteX SoC (vexiiriscv.soc.litex.SocGen)--with-swddebug_swd_swd_swclk, debug_swd_swd_swdio_{read,write,writeEnable}
MicroSoc (vexiiriscv.soc.micro.MicroSocGen)--jtag-tap=false --swd=truesocCtrl_debugModule_swd_swd_*
  • One DTM at a time (RISC-V Debug Spec Ch. 6): --with-swd together with --with-jtag-tap / --with-jtag-instruction is refused at elaboration.
  • SWDIO is three wires (read / write / writeEnable); the tristate belongs to the integrator.
  • A --with-jtag-tap netlist is unchanged by #184.

LiteX integration — pending. LiteX's cpu/vexiiriscv still pins a VexiiRiscv revision without SWD, and its --with-swd-debug option for VexiiRiscv (four ports, add_swd(), reset wiring) is not published yet; it will be proposed together with the pin bump.

Hardware results

Same MCU-Link, wiring and OpenOCD fork as the VexRiscv lane above.

ConfigurationBoardBuildResult
RV32 linux (RV32IMA + S/U), 1 hartArty A7-35T, 100 MHzWNS 0.233 ns, 41.6 % LUTDPIDR 0x0ba11aab, AP_IDR 0x74726976, misa=0x40141101; halt / step / resume; GDB load (67 KB/s), break main / help, stepi, bt
RV64 debian (RV64IMAFDC + S/U), 1 hartArty A7-35T, 100 MHzWNS 0.131 ns, 75 % LUT / 86 % BRAMXLEN=64, misa=0x800000000014112d; halt / step / resume; GDB (set arch riscv:rv64): load, breakpoints, 64-bit register write / read-back, FPU registers (fcsr, ft0), bt
RV64 debian, 2 hartsArty A7-100T, 80 MHzWNS 0.532 ns, 44 % LUTboth harts on one DTM; independent halt / step / resume per hart; GDB sees one thread per hart

Two harts of the RV64 variant do not fit the A7-35T (estimated ~128 % LUT, ~126 % BRAM), hence A7-100T at 80 MHz is used which meets timing.

Log lines that are expected on VexiiRiscv:

  • Found 0 triggers — LiteX's default VexiiRiscv configuration has no hardware triggers (same over JTAG); software breakpoints in RAM work.
  • Failed to read memory (addr=0x3ffffffc) — GDB peeks at the word before _start, which is unmapped; the DM correctly reports the bus error.
  • Core N could not be made part of halt group 1 (two harts) — this DM implements no halt groups, which the spec allows; OpenOCD halts the harts one after the other.

Two harts: OpenOCD configuration

One riscv target per hart, all on the same DTM (DMI is shared by every hart behind the DM; the hart is chosen by -coreid, i.e. hartsel). Source it in place of vexriscv_swd_riscv_master.cfg, after openocd_arty_swd.cfg and vexriscv_swd.cfg.

For GDB, one thread per hart:

dtm create vexriscv.dtm -type vexriscv-gateway -dap vexriscv.dap -ap-num 0

target create vexriscv.rv0 riscv -dtm vexriscv.dtm -coreid 0 -rtos hwthread
target create vexriscv.rv1 riscv -dtm vexriscv.dtm -coreid 1 -rtos hwthread
target smp vexriscv.rv0 vexriscv.rv1

riscv set_command_timeout_sec 120
vexriscv.rv0 configure -event gdb-attach halt
set arch riscv:rv64
target extended-remote localhost:3333
info threads
#  * 1  Thread 1 "vexriscv.rv0" ...
#    2  Thread 2 "vexriscv.rv1" ... 0x00000000000000a4 in ?? ()
thread 2
p/x $mhartid        # 0x1

For independent per-hart control, drop -rtos hwthread and target smp, then select a hart with targets vexriscv.rv0 / targets vexriscv.rv1. Halting hart 1 leaves hart 0 running (vexriscv.rv0 curstate → running, vexriscv.rv1 curstate → halted), and each hart steps and resumes on its own.


DMI gateway vs Mem-AP (the RP2350 approach)

The closest production precedent is the Raspberry Pi RP2350: SWD pins, an ARM SW-DP, and a RISC-V Debug Module behind it (RP2350 datasheet §3.5, §3.8). Both designs are the same up to one layer. They differ only in the access port between the DP and the DM:

RP2350This design (VexRiscv / VexiiRiscv)
AP typestandard CoreSight APB Mem-AP, at 0x0a000 in the debug address spacecustom 4-register designer AP — AP_IDR / DMI_ADDR / DMI_DATA / POSTED_READ (Phase 2C)
Reaching DM register nmemory-mapped at n × 4: write the address to TAR, then access DRW (or BD0–BD3)write n to DMI_ADDR, then read / write DMI_DATA
SWD packets per DM access1–4, depending on the access pattern1–2, depending on the access pattern
DP architectureADIv6 (SELECT holds the AP base address)ADIv5 (APSEL field)
DP state at power-upDormant (needs the selection-alert wake-up sequence)active
Standing in the RISC-V Debug Speccustom DTM (Ch. 6)custom DTM (Ch. 6) — the same

Why this design keeps the gateway. What a Mem-AP would buy is abillity to speak Mem-AP lingo from ARM world which on the wire it is not cheap.

Speed

Both designs are an address register plus a data register — TAR + DRW on a Mem-AP, DMI_ADDR + DMI_DATA on the gateway However, the real difference is the Mem-AP's banked window. Host side tooling like OpenOCD's Mem-AP layer reads through BD0–BD3: TAR is set to a 16-byte-aligned address, and the four words of that window are then reached without touching TAR again. But BD0–BD3 sit in a different AP register bank from TAR, so moving to a new window costs a DP SELECT write, the TAR write, and a SELECT write back.

The gateway keeps all its registers in bank 0 and never writes SELECT.

This is explained below for SWD packets per DM access for both the two backends:

AccessGatewayMem-AP (BD path)
Same register as the previous access11
Another register in the same 4-register window21
Register in a different window24 (SELECT + TAR + SELECT + BD)
End of a batch that read something+ 1 (RDBUFF)+ 1 (RDBUFF)

Also RISC-V DM's registers cluster in those windows — dmcontrol / dmstatus, abstractcs / command, and data0–data3 each share one — so for real operations:

OperationGatewayMem-AP
Halt: write dmcontrol, poll dmstatus N times2 + 2 + (N − 1)4 + N
Read a GPR, RV32: write command, poll abstractcs, read data0710
Read a GPR, RV64: also read data1911
Poll the same register1 each1 each

So on the wire. the gateway is a few packets ahead.

Area

In area the gateway is clearly smaller. Its AP is a 7-bit address register plus one DebugBus request path; the whole SWD DTM (PHY, DP, gateway, clock crossing) is 216 LUTs on an Artix-7. The SWD DTM breaks down as:

Part of DebugTransportModuleSwdLUTsFFsAlso needed with a Mem-AP?
SwdPhy — wire protocol95130yes, same DP
SwdDp — DP registers2550yes
Response clock crossing (FlowCCByToggle)3470yes — any AP needs a crossing to the DM (Hazard3 uses an async APB bridge)
Gateway AP + command clock crossing60115no — this is the part a Mem-AP would replace
Total216409

So the gateway costs at most 60 LUTs / 115 FFs, and part of that is the command-side clock crossing a Mem-AP would need too. A Mem-AP in its place needs at least a 32-bit TAR, a CSW register (access size, auto-increment, protection), the banked BD decode, an address incrementer and an APB manager — roughly 100–150 LUTs by estimate (not synthesised).

Context - Mixed Architecture vs Pure RISC-V

RP2350 is a dual-architecture part: each core slot holds an Arm Cortex-M33 and a Hazard3, selected at boot. Its debug complex is Arm CoreSight around one SW-DP, with two AHB5 Mem-APs (debug address space 0x02000 / 0x04000) for the two Cortex-M33s, the APB Mem-AP at 0x0a000 for the RISC-V DM, and RP-AP for chip-level control (datasheet §3.5.2–3.5.3, Figure 6). The SW-DP and the Mem-AP infrastructure are there for the Arm cores anyway. And the Hazard3 DM already exposes a byte-addressed APB port, as shown above. Connecting it through one more APB Mem-AP therefore costs almost nothing extra, needs no new transport RTL, and makes the RISC-V cores reachable through the same port, the same way, as the Arm ones.

This design starts from the opposite position: there is no Arm debug IP to reuse. The SW-DP is written from scratch in SpinalHDL, and the SpinalHDL DM exposes a word-addressed DebugBus(7), not APB. Building a Mem-AP would add a CoreSight-shaped register set and an APB manager only to reach a DM that doesn't speak APB; the gateway reaches DebugBus directly. The one thing the Mem-AP shape would buy is support in tools that only speak Mem-AP — and that is a host-side problem, solved here by the vexriscv-gateway OpenOCD backend without touching the RTL.

Mem-AP fits whenGateway fits when
Debug IP already on chipan Arm CoreSight DAP is present (e.g. for Arm cores on the same die)the DP is your own RTL
DM portAPB, byte-addressed (Hazard3)word-addressed bus (DebugBus)
Host toolingmust work with tools that only drive Mem-APsyou control the host side (OpenOCD backend)
Areathe AP is already paid forthe smallest AP that does the job

RISC-V-only chips: why the gateway is usually the better choice

Take the Arm cores away and RP2350's main reason for a Mem-AP disappears: there is no CoreSight debug IP on the die to reuse. For a SoC or MCU whose only processors are RISC-V, the gateway is usually the better fit:

  • No CoreSight IP to reuse. The SW-DP has to be built anyway (here it is SpinalHDL). A Mem-AP on top of it adds TAR, CSW, the banked BD decode and an APB manager, with no functional gain: the DM is the only thing behind the AP.
  • The DM has a simple word-addressed port. SpinalHDL's DebugModule speaks DebugBus(7); the gateway connects to it directly, whereas a Mem-AP would first need an APB front end on the DM.
  • It is the smaller AP, and on the wire it is even or slightly ahead (see the tables above) — by tens of LUTs and a few packets per operation, so a tie-breaker rather than the main reason.

CPU-embedded debug plugin (EmbeddedRiscvJtag) over SWD

Both CPUs also have a plugin, EmbeddedRiscvJtag, that builds the Debug Module and its transport inside a single-hart CPU. It is meant for small SoCs that instantiate the CPU directly. Examples are VexiiRiscv's standalone Generate (--debug-jtag-*), VexRiscv's GenFullWithOfficialRiscvDebug and Briey, and third-party SoCs such as aesc-silicon's nafarr / ElemRV. The plugin's SWD mode reuses the same DebugTransportModuleSwd as the SMP cluster and the SoC-level debug fiber:

PROptionHardware check
VexiiRiscv #188 — merged 2026-09-28 (19b41a7)withSwd on the plugin, --debug-swdArty A7-100T, RV32, Black Magic Probe: load, software and single-step lane identical to the MCU-Link values
VexRiscv #500 — merged 2026-09-27 (aefc0e0)withSwd on the pluginArty A7-35T, LiteX "standard" core + official debug: MCU-Link (65 KB/s load, software + hardware breakpoints, ndmreset) and Black Magic Probe, identical results

On both merges the existing JTAG configurations generate netlists identical to dev, and the SWD blocks are byte-identical to the SMP cluster's. The host side (probe, OpenOCD fork, configs) is unchanged. The plugin supports one hart; multi-hart designs use the SoC-level paths above.

ElemRV

elements-nafarr#72 adds SWD debug transport to ElemRV using EmbeddedRiscvJtag.

elements-zibal#62 passes it through the zibal platforms: a board sets debugTransport = DebugTransport.Swd in Hydrogen.Parameter (or Carbon / Nitrogen) and wires two pads, swclk and swdio, instead of four JTAG pads.


Black Magic Probe (OpenOCD alternative)

A Black Magic Probe (BMP) runs the GDB server on the probe: GDB connects straight to its USB serial port, with no OpenOCD in between. Stock BMP firmware reads this design's SW-DP but cannot use the DMI gateway AP. Support is proposed in blackmagic #2322 (riscv_adi_dtm: support the SpinalHDL SWD "DMI gateway" AP; branch disdi:feature/riscv-swd-dmi-gateway). No RTL change was needed.

Wiring (BMP v2.3, same Arty Pmod JB harness)

BMP 10-pinSignalArty
1VTref (sets the probe's I/O level)JB12 (3.3 V)
2SWDIOJB7
4SWCLKJB3
3 / 5 / 9GNDJB11, plus the second ground JP2.3 → JB5

Hardware results

Board / imageResult
Arty A7-35T, VexRiscv SMP (RV32IMA), golden imagerv32ima (exts 00141101). load demo.elf (6552 B), break main 0x4000069c, stepi → 0x400006b0, break help 0x40000644, bt #0 help / #1 0x400006cc main — identical to the OpenOCD + MCU-Link lane; compare-sections matches
Arty A7-100T, VexiiRiscv RV64 × 2 hartsDM v0.13, both harts found as two targets; attach, registers, single step, resume/halt per hart

Use

# Probe firmware (BMP v2.x), from the blackmagic tree with #2322  :
meson setup build-fw --cross-file cross-file/bmp-v1-v2-riscv.ini && ninja -C build-fw
dfu-util -d 1d50:6018,:6017 -s 0x08002000:leave -D build-fw/blackmagic_bmp_v1_v2_firmware.bin

# GDB, straight to the probe (no OpenOCD):
riscv64-unknown-elf-gdb -ex 'set arch riscv:rv32' -ex 'set mem inaccessible-by-default off' \
  -ex 'target extended-remote /dev/ttyACM0' -ex 'monitor swdp_scan' -ex 'attach 1' demo/demo.elf
(gdb) load
(gdb) break main
(gdb) continue

Testing

Phase 4: Verification and Testing

Host-visible sim paths that already pass are tracked under Phase 3 (raw-AP smoke, riscv examine, GDB attach/regs, demo break/continue). This chapter is the broader regression matrix — much of it still open, especially on hardware.


Done in sim (via Phase 3)

  • SWD-DTM → DebugBus → DM path at sbt (Phases 2A–2C) and LiteX Verilator SoC
  • OpenOCD master smoke: DPIDR + dap apreg → dmstatus
  • OpenOCD riscv examine + halt/resume/register access over SWD (patched builds)
  • GDB attach / registers / memory R/W over SWD
  • GDB break main / continue / backtrace on preloaded demo (no GDB load)

Still open

  • Automated SWD stimuli via custom OpenOCD target on real CMSIS-DAP hardware (Arty)
  • Test dmactive activation/deactivation sequence (formal matrix)
  • Test abstract command error handling (all 7 cmderr codes)
  • Test EBREAK behavior with dcsr.ebreakm / ebreaks / ebreaku combinations
  • Test single-step across privilege mode transitions
  • Test trigger module: each trigger type, chaining, dmode security
  • Test System Bus Access error handling (if implemented)
  • Test authentication mechanism (if implemented)
  • Decide dmstatus.version claim (2 = 0.13 vs 3 = 1.0)
  • Regression suite for all implemented features (CI-friendly)

Documentation

Phase 5: Documentation & Developer Interface


Done

  • Phase 2A–2C architecture chapters in this book (SwdPhy / SwdDp / SwdDmiGateway)
  • Operator how-to for sim three-terminal attach (JTAG and SWD) — Phase 3
  • Operator how-to for Arty hardware (JTAG via BSCANE2 USER4; SWD via MCU-Link CMSIS-DAP on Pmod JB) — Phase 3
  • Document placeholder DPIDR, AP_IDR, and Phase 2C DMI_ADDR / DMI_DATA / POSTED_READ map (see Phase 2C)
  • Published OpenOCD Vexriscv fork: disdi/openocd vexriscv-gateway (9786 + fixes + gateway; still unmerged upstream)

Still open

  • Broader user-facing guide beyond this stack (probe-rs, multi-probe notes)
  • Finalize published identification constants once IDs leave placeholder status
  • Document which optional RISC-V debug-spec features are implemented vs stubbed (dmstatus.version, abstract commands surface, triggers, SBA)