An in-progress, educational implementation of a 32-bit RISC-V processor in
SystemVerilog. The project is working toward a five-stage RV32I pipeline with
data forwarding, hazard handling, branch/jump control, separate instruction and
data memories, and a verification environment.
Important
This repository is not a complete CPU yet. It currently contains several datapath building blocks and a partial instruction-fetch stage. There is no top-level processor, decoder/control unit, pipeline-register set, hazard unit, forwarding unit, or testbench. Some source files also need the syntax and integration fixes listed in Known issues before the RTL can compile as a complete design.
The diagram is the architectural target for the project. Instructions flow through these stages:
- IF — Instruction Fetch: select the next PC, increment the PC by four, and fetch an instruction from instruction memory.
- ID — Instruction Decode / Register Read: decode the instruction, read
rs1andrs2, generate the immediate, and create control signals. - EX — Execute: select forwarded operands, perform the ALU operation, calculate branch/jump targets, and evaluate branch conditions.
- MEM — Memory Access: read or write data memory for load/store instructions.
- WB — Write Back: select the ALU result, load data,
PC + 4, or another required value and write it tord.
Pipeline registers are intended between IF/ID, ID/EX, EX/MEM, and MEM/WB. The forwarding paths shown in the diagram are intended to resolve most read-after- write dependencies without stalling. A hazard-detection unit will still be needed for cases such as load-use dependencies, and taken control transfers must flush younger instructions.
| Area | Status | Current capability |
|---|---|---|
| Parameterized adder | Partial | Combinational in + ADDEND |
| Generic multiplexer | Partial | Selects one of NUM_INPUT packed inputs |
| Program counter | Implemented as a block | 32-bit register with asynchronous active-low reset |
| Instruction memory | Partial | 1,024 × 32-bit word array loaded from imem.hex |
| Instruction-fetch stage | Partial | PC selection, PC + 4, instruction fetch, registered instruction output |
| Register file | Partial | 32 registers, two asynchronous reads, one synchronous write, reset-to-zero |
| Immediate generator | Partial | I, S, U, J, and B immediate formats |
| ALU | Partial | Add, subtract, bitwise logic, shifts, signed/unsigned set-less-than |
| Branch comparator | Partial | BEQ, BNE, BLT, BGE, BLTU, and BGEU decisions |
| Data memory | Partial | 1,024 × 32-bit words with synchronous word write and asynchronous word read |
| Decoder/control unit | Not started | Opcode/funct decoding and control generation are missing |
| Pipeline registers | Not started | IF/ID, ID/EX, EX/MEM, and MEM/WB registers are missing |
| Forwarding unit | Not started | Forwarding selection and datapaths are missing |
| Hazard/flush unit | Not started | Stall, bubble, and flush logic are missing |
| Top-level core | Not started | Existing blocks are not connected into a processor |
| Verification | Not started | No testbench, assertions, reference model, or regression exists |
| FPGA/synthesis flow | Not started | No constraints, board wrapper, or build scripts exist |
“Partial” means that a useful block exists, but it has not yet been fully integrated and verified against the ISA.
.
├── images/
│ └── forwarding_datapath.png # Target pipelined datapath
├── rtl/
│ ├── IF.sv # Partial instruction-fetch stage
│ ├── adder.sv # Constant-add combinational block
│ ├── alu.sv # Integer ALU
│ ├── bramch_logic.sv # Branch comparator (filename typo retained)
│ ├── dmem.sv # Word-addressed data memory
│ ├── imem.sv # Word-addressed instruction memory
│ ├── imm_gen.sv # Immediate extraction/sign extension
│ ├── mux.sv # Parameterized multiplexer
│ ├── pc.sv # Program-counter register
│ └── register_file.sv # 32 × 32-bit integer register file
└── README.md
Adds the constant parameter ADDEND to in. The fetch stage uses it to form
PC + 4.
A combinational packed-array multiplexer. sel has
$clog2(NUM_INPUT) bits; callers must ensure sel never addresses beyond the
configured number of inputs.
Updates the program counter on each rising clock edge and resets it to address
zero when rst_n is low. The reset is asynchronous and active-low.
Models 4 KiB of instruction storage as 1,024 32-bit words. It indexes the
memory with addr[11:2], so instructions are expected to be 4-byte aligned.
Simulation loads words with $readmemh("imem.hex", mem).
Combines the PC register, constant adder, next-PC multiplexer, and instruction
memory. pc_sel = 0 chooses PC + 4; pc_sel = 1 chooses pc_branch.
The fetched instruction is registered on the rising clock edge.
Implements the RV32 integer register bank with two combinational read ports and
one rising-edge write port. Writes to register x0 are suppressed. All 32
entries are cleared while the asynchronous active-low resN reset is asserted.
Generates sign-extended I-, S-, J-, and B-type immediates and the shifted
U-type immediate. Its current imm_sel encoding is:
imm_sel |
Format |
|---|---|
3'b000 |
I-type |
3'b001 |
S-type |
3'b010 |
U-type |
3'b011 |
J-type |
3'b100 |
B-type |
The ALU uses the following internal control encoding. This is an implementation detail, not a RISC-V instruction encoding.
alu_control |
Operation | Typical instructions |
|---|---|---|
4'b0000 |
Add | ADD, ADDI, address calculation |
4'b0001 |
Subtract | SUB |
4'b0010 |
Bitwise AND | AND, ANDI |
4'b0011 |
Bitwise OR | OR, ORI |
4'b0100 |
Bitwise XOR | XOR, XORI |
4'b0101 |
Logical left shift | SLL, SLLI |
4'b0110 |
Logical right shift | SRL, SRLI |
4'b0111 |
Arithmetic right shift | SRA, SRAI |
4'b1000 |
Signed less-than | SLT, SLTI |
4'b1001 |
Unsigned less-than | SLTU, SLTIU |
zero_flag is asserted whenever the selected result is zero.
Compares two 32-bit operands and produces pc_sel. Its current br_type
encoding is:
br_type |
Meaning |
|---|---|
3'b000 |
No branch |
3'b001 |
BEQ |
3'b010 |
BNE |
3'b011 |
BLT |
3'b100 |
BGE |
3'b101 |
BLTU |
3'b110 |
BGEU |
Models 4 KiB of data storage as 1,024 32-bit words. A write occurs on a rising
clock edge when we is high; reads are combinational. The existing block only
implements full-word accesses and ignores addr[1:0].
The intended baseline is the unprivileged RV32I integer ISA. The table below
describes the goal, not the current end-to-end capability.
| Class | Instructions |
|---|---|
| Register arithmetic/logic | ADD, SUB, SLL, SLT, SLTU, XOR, SRL, SRA, OR, AND |
| Immediate arithmetic/logic | ADDI, SLTI, SLTIU, XORI, ORI, ANDI, SLLI, SRLI, SRAI |
| Loads | LB, LH, LW, LBU, LHU |
| Stores | SB, SH, SW |
| Conditional branches | BEQ, BNE, BLT, BGE, BLTU, BGEU |
| Jumps | JAL, JALR |
| Upper immediates | LUI, AUIPC |
FENCE, ECALL, EBREAK, CSRs, traps/interrupts, privilege modes, and ISA
extensions such as M, A, C, and floating point are outside the initial
milestone unless deliberately added later.
- Both memories contain 1,024 words and therefore cover 4 KiB each.
addr[11:2]selects a word; address bits above bit 11 are currently ignored.- Instruction memory is read-only from the core and has an asynchronous model.
- Data-memory reads are asynchronous and writes are synchronous.
- The current data memory has no byte-enable logic, so byte and halfword load/store behavior is not implemented.
- There is no memory-mapped I/O, bus interface, alignment exception, or access fault handling yet.
For simulation, imem.hex must contain one 32-bit hexadecimal instruction per
line in the order expected by $readmemh, for example:
00500093
00308113
002081b3
The repository does not currently include an imem.hex program.
pcandif_stageuse active-low reset signalrst_n.register_fileuses the differently named active-low resetresN.- The PC and register file reset asynchronously.
- Register-file and data-memory writes occur on the rising clock edge.
- Instruction and data memory reads are modeled as combinational reads.
These conventions should be made consistent before integration. If the design targets FPGA block RAM, memory timing may also need to become synchronous, with the pipeline adjusted accordingly.
No simulator, file list, build script, or testbench is included yet, so there is no working project-level build command at this revision. After completing Roadmap Phase 1, a minimal Icarus Verilog flow could look like:
iverilog -g2012 -s tb_rv32i_core -o build/core.vvp rtl/*.sv tb/*.sv
vvp build/core.vvpAn equivalent Verilator lint command could be:
verilator --lint-only --Wall --top-module rv32i_core rtl/*.svThese commands are targets for the repository’s future build setup; they will not succeed against the current revision because the top-level core and testbench do not exist and the parameter declarations still need correction.
Verification should be added incrementally instead of waiting for the complete pipeline:
- Write self-checking unit tests for the ALU, immediate generator, register file, branch comparator, PC, and memories.
- Add assertions for
x0 == 0, legal control selections, aligned instruction fetches, stable stalled pipeline registers, and correct flush behavior. - Test each instruction with directed assembly programs, including negative immediates and signed/unsigned boundary values.
- Add dependency tests for EX-to-EX, MEM-to-EX, and WB-to-ID forwarding,
load-use stalls, consecutive hazards, and writes to
x0. - Add control-flow tests for taken/not-taken branches, back-to-back branches, JAL/JALR link values, and wrong-path flushing.
- Compare architectural state against a trusted RISC-V reference model and run a standard RV32I architectural test suite.
- Add lint and regression jobs to continuous integration, then collect code and functional coverage.
- Replace semicolons with commas in parameter-port lists.
- Rename
bramch_logic.svtobranch_logic.sv. - Standardize reset naming and reset behavior.
- Define shared constants/types for ALU, immediate, branch, and write-back selections instead of scattering raw bit encodings.
- Decide whether the IF stage exposes a combinational instruction or a
matched
(PC, instruction)IF/ID pair. - Add default initialization or explicit unknown-state handling for memory.
- Add a simulator file list and a lint configuration.
Exit criterion: every existing RTL block lints and compiles, with no top- level integration assumed.
- Create
tb/and self-checking tests for every existing module. - Cover ALU overflow wraparound, signed comparisons, all shift amounts,
branch boundaries, all immediate layouts,
x0, and memory addressing. - Add a reproducible
Makefileor script for lint, unit tests, and waveform generation. - Commit small known-good
.hexprograms or generate them during tests.
Exit criterion: all blocks pass automated unit tests and lint locally.
- Implement instruction decoding and the control unit.
- Add operand, write-back, and next-PC selection logic.
- Add byte enables plus load extraction/sign extension for LB/LH/LW/LBU/LHU and SB/SH/SW.
- Implement LUI, AUIPC, JAL, and JALR.
- Connect the existing blocks under an
rv32i_coretop level. - Verify all planned instructions before introducing pipeline hazards.
Exit criterion: directed programs execute correctly in a complete non- pipelined core.
- Add valid bits and IF/ID, ID/EX, EX/MEM, and MEM/WB registers.
- Partition datapath and control signals by stage.
- Carry the PC and
PC + 4to the stages that need them. - Confirm that retirement and memory side effects occur exactly once.
Exit criterion: hazard-free programs match the single-cycle reference.
- Implement EX/MEM and MEM/WB forwarding for both ALU operands.
- Forward store data and operands used by branch comparison where needed.
- Detect load-use hazards and insert a one-cycle bubble.
- Define branch/jump resolution stage and flush all younger wrong-path work.
- Handle simultaneous stall, flush, and forwarding conditions explicitly.
Exit criterion: dependency-heavy and control-flow tests pass without manual NOP insertion.
- Add illegal-instruction handling or document the chosen behavior.
- Decide and document misaligned-access behavior.
- Run RV32I architectural compliance tests and differential tests.
- Add SystemVerilog assertions and functional coverage for hazards and instruction classes.
- Add continuous-integration lint and regression jobs.
Exit criterion: the defined RV32I subset passes its automated compliance and regression suite.
- Add a synthesis-friendly memory or external instruction/data bus.
- Choose an FPGA board, add clock/reset wrappers and timing constraints, and check timing/resource reports.
- Add memory-mapped UART/GPIO and a small software demo if desired.
- Measure maximum frequency, CPI, and resource use.
- Consider traps/interrupts, CSRs,
Zicsr,M, or caches only after the baseline core is stable.
Exit criterion: a documented FPGA or ASIC-oriented build runs a repeatable demo and meets its stated timing target.
- Several parameterized modules currently terminate parameter declarations with semicolons instead of separating them with commas; this is invalid SystemVerilog syntax.
bramch_logic.svis misspelled, although its module is namedbranch_logic.- The registered
instructionand combinationalpc_outinif_stagemay refer to different fetch cycles. A proper IF/ID register should keep them together. if_stagehard-codes 32-bitpcandimeminstances even though the stage exposes aDATA_WIDTHparameter.- Reset signal names are inconsistent (
rst_nandresN). imemrequires an absent working-directory-relativeimem.hexfile.- Memories have no reset/initial value policy, bounds checking, byte enables, or alignment checking.
dmemonly supports word operations, so the full RV32I load/store set is not yet implemented.- The register file clears all entries in one asynchronous-reset loop. This is convenient for simulation but may not map efficiently to FPGA memory.
- No block validates its raw control encoding, and there is no shared package to prevent encoding mismatches.
- There is no license file yet; add one before defining reuse/distribution terms.
Keep each change small and verifiable: add or update a test with every RTL change, run lint and the full regression, and document any new interface or control encoding. Completion should be judged by the exit criteria above, not only by the presence of modules.
