Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

SystemVerilog RV32I Pipelined Processor

An in-progress, educational implementation of a 32-bit RISC-V processor in SystemVerilog. The project is working toward a five-stage RV32I pipeline with data forwarding, hazard handling, branch/jump control, separate instruction and data memories, and a verification environment.

Important

This repository is not a complete CPU yet. It currently contains several datapath building blocks and a partial instruction-fetch stage. There is no top-level processor, decoder/control unit, pipeline-register set, hazard unit, forwarding unit, or testbench. Some source files also need the syntax and integration fixes listed in Known issues before the RTL can compile as a complete design.

Target architecture

Target five-stage RV32I datapath

The diagram is the architectural target for the project. Instructions flow through these stages:

  1. IF — Instruction Fetch: select the next PC, increment the PC by four, and fetch an instruction from instruction memory.
  2. ID — Instruction Decode / Register Read: decode the instruction, read rs1 and rs2, generate the immediate, and create control signals.
  3. EX — Execute: select forwarded operands, perform the ALU operation, calculate branch/jump targets, and evaluate branch conditions.
  4. MEM — Memory Access: read or write data memory for load/store instructions.
  5. WB — Write Back: select the ALU result, load data, PC + 4, or another required value and write it to rd.

Pipeline registers are intended between IF/ID, ID/EX, EX/MEM, and MEM/WB. The forwarding paths shown in the diagram are intended to resolve most read-after- write dependencies without stalling. A hazard-detection unit will still be needed for cases such as load-use dependencies, and taken control transfers must flush younger instructions.

Current implementation status

Area Status Current capability
Parameterized adder Partial Combinational in + ADDEND
Generic multiplexer Partial Selects one of NUM_INPUT packed inputs
Program counter Implemented as a block 32-bit register with asynchronous active-low reset
Instruction memory Partial 1,024 × 32-bit word array loaded from imem.hex
Instruction-fetch stage Partial PC selection, PC + 4, instruction fetch, registered instruction output
Register file Partial 32 registers, two asynchronous reads, one synchronous write, reset-to-zero
Immediate generator Partial I, S, U, J, and B immediate formats
ALU Partial Add, subtract, bitwise logic, shifts, signed/unsigned set-less-than
Branch comparator Partial BEQ, BNE, BLT, BGE, BLTU, and BGEU decisions
Data memory Partial 1,024 × 32-bit words with synchronous word write and asynchronous word read
Decoder/control unit Not started Opcode/funct decoding and control generation are missing
Pipeline registers Not started IF/ID, ID/EX, EX/MEM, and MEM/WB registers are missing
Forwarding unit Not started Forwarding selection and datapaths are missing
Hazard/flush unit Not started Stall, bubble, and flush logic are missing
Top-level core Not started Existing blocks are not connected into a processor
Verification Not started No testbench, assertions, reference model, or regression exists
FPGA/synthesis flow Not started No constraints, board wrapper, or build scripts exist

“Partial” means that a useful block exists, but it has not yet been fully integrated and verified against the ISA.

Repository layout

.
├── images/
│   └── forwarding_datapath.png  # Target pipelined datapath
├── rtl/
│   ├── IF.sv                    # Partial instruction-fetch stage
│   ├── adder.sv                 # Constant-add combinational block
│   ├── alu.sv                   # Integer ALU
│   ├── bramch_logic.sv          # Branch comparator (filename typo retained)
│   ├── dmem.sv                  # Word-addressed data memory
│   ├── imem.sv                  # Word-addressed instruction memory
│   ├── imm_gen.sv               # Immediate extraction/sign extension
│   ├── mux.sv                   # Parameterized multiplexer
│   ├── pc.sv                    # Program-counter register
│   └── register_file.sv         # 32 × 32-bit integer register file
└── README.md

Existing RTL blocks

adder

Adds the constant parameter ADDEND to in. The fetch stage uses it to form PC + 4.

mux

A combinational packed-array multiplexer. sel has $clog2(NUM_INPUT) bits; callers must ensure sel never addresses beyond the configured number of inputs.

pc

Updates the program counter on each rising clock edge and resets it to address zero when rst_n is low. The reset is asynchronous and active-low.

imem

Models 4 KiB of instruction storage as 1,024 32-bit words. It indexes the memory with addr[11:2], so instructions are expected to be 4-byte aligned. Simulation loads words with $readmemh("imem.hex", mem).

if_stage

Combines the PC register, constant adder, next-PC multiplexer, and instruction memory. pc_sel = 0 chooses PC + 4; pc_sel = 1 chooses pc_branch. The fetched instruction is registered on the rising clock edge.

register_file

Implements the RV32 integer register bank with two combinational read ports and one rising-edge write port. Writes to register x0 are suppressed. All 32 entries are cleared while the asynchronous active-low resN reset is asserted.

imm_gen

Generates sign-extended I-, S-, J-, and B-type immediates and the shifted U-type immediate. Its current imm_sel encoding is:

imm_sel Format
3'b000 I-type
3'b001 S-type
3'b010 U-type
3'b011 J-type
3'b100 B-type

alu

The ALU uses the following internal control encoding. This is an implementation detail, not a RISC-V instruction encoding.

alu_control Operation Typical instructions
4'b0000 Add ADD, ADDI, address calculation
4'b0001 Subtract SUB
4'b0010 Bitwise AND AND, ANDI
4'b0011 Bitwise OR OR, ORI
4'b0100 Bitwise XOR XOR, XORI
4'b0101 Logical left shift SLL, SLLI
4'b0110 Logical right shift SRL, SRLI
4'b0111 Arithmetic right shift SRA, SRAI
4'b1000 Signed less-than SLT, SLTI
4'b1001 Unsigned less-than SLTU, SLTIU

zero_flag is asserted whenever the selected result is zero.

branch_logic

Compares two 32-bit operands and produces pc_sel. Its current br_type encoding is:

br_type Meaning
3'b000 No branch
3'b001 BEQ
3'b010 BNE
3'b011 BLT
3'b100 BGE
3'b101 BLTU
3'b110 BGEU

dmem

Models 4 KiB of data storage as 1,024 32-bit words. A write occurs on a rising clock edge when we is high; reads are combinational. The existing block only implements full-word accesses and ignores addr[1:0].

Planned instruction support

The intended baseline is the unprivileged RV32I integer ISA. The table below describes the goal, not the current end-to-end capability.

Class Instructions
Register arithmetic/logic ADD, SUB, SLL, SLT, SLTU, XOR, SRL, SRA, OR, AND
Immediate arithmetic/logic ADDI, SLTI, SLTIU, XORI, ORI, ANDI, SLLI, SRLI, SRAI
Loads LB, LH, LW, LBU, LHU
Stores SB, SH, SW
Conditional branches BEQ, BNE, BLT, BGE, BLTU, BGEU
Jumps JAL, JALR
Upper immediates LUI, AUIPC

FENCE, ECALL, EBREAK, CSRs, traps/interrupts, privilege modes, and ISA extensions such as M, A, C, and floating point are outside the initial milestone unless deliberately added later.

Memory model

  • Both memories contain 1,024 words and therefore cover 4 KiB each.
  • addr[11:2] selects a word; address bits above bit 11 are currently ignored.
  • Instruction memory is read-only from the core and has an asynchronous model.
  • Data-memory reads are asynchronous and writes are synchronous.
  • The current data memory has no byte-enable logic, so byte and halfword load/store behavior is not implemented.
  • There is no memory-mapped I/O, bus interface, alignment exception, or access fault handling yet.

For simulation, imem.hex must contain one 32-bit hexadecimal instruction per line in the order expected by $readmemh, for example:

00500093
00308113
002081b3

The repository does not currently include an imem.hex program.

Reset and timing conventions

  • pc and if_stage use active-low reset signal rst_n.
  • register_file uses the differently named active-low reset resN.
  • The PC and register file reset asynchronously.
  • Register-file and data-memory writes occur on the rising clock edge.
  • Instruction and data memory reads are modeled as combinational reads.

These conventions should be made consistent before integration. If the design targets FPGA block RAM, memory timing may also need to become synchronous, with the pipeline adjusted accordingly.

Building and simulation

No simulator, file list, build script, or testbench is included yet, so there is no working project-level build command at this revision. After completing Roadmap Phase 1, a minimal Icarus Verilog flow could look like:

iverilog -g2012 -s tb_rv32i_core -o build/core.vvp rtl/*.sv tb/*.sv
vvp build/core.vvp

An equivalent Verilator lint command could be:

verilator --lint-only --Wall --top-module rv32i_core rtl/*.sv

These commands are targets for the repository’s future build setup; they will not succeed against the current revision because the top-level core and testbench do not exist and the parameter declarations still need correction.

Verification plan

Verification should be added incrementally instead of waiting for the complete pipeline:

  1. Write self-checking unit tests for the ALU, immediate generator, register file, branch comparator, PC, and memories.
  2. Add assertions for x0 == 0, legal control selections, aligned instruction fetches, stable stalled pipeline registers, and correct flush behavior.
  3. Test each instruction with directed assembly programs, including negative immediates and signed/unsigned boundary values.
  4. Add dependency tests for EX-to-EX, MEM-to-EX, and WB-to-ID forwarding, load-use stalls, consecutive hazards, and writes to x0.
  5. Add control-flow tests for taken/not-taken branches, back-to-back branches, JAL/JALR link values, and wrong-path flushing.
  6. Compare architectural state against a trusted RISC-V reference model and run a standard RV32I architectural test suite.
  7. Add lint and regression jobs to continuous integration, then collect code and functional coverage.

Roadmap

Phase 0 — Make the existing blocks clean and compilable

  • Replace semicolons with commas in parameter-port lists.
  • Rename bramch_logic.sv to branch_logic.sv.
  • Standardize reset naming and reset behavior.
  • Define shared constants/types for ALU, immediate, branch, and write-back selections instead of scattering raw bit encodings.
  • Decide whether the IF stage exposes a combinational instruction or a matched (PC, instruction) IF/ID pair.
  • Add default initialization or explicit unknown-state handling for memory.
  • Add a simulator file list and a lint configuration.

Exit criterion: every existing RTL block lints and compiles, with no top- level integration assumed.

Phase 1 — Unit verification

  • Create tb/ and self-checking tests for every existing module.
  • Cover ALU overflow wraparound, signed comparisons, all shift amounts, branch boundaries, all immediate layouts, x0, and memory addressing.
  • Add a reproducible Makefile or script for lint, unit tests, and waveform generation.
  • Commit small known-good .hex programs or generate them during tests.

Exit criterion: all blocks pass automated unit tests and lint locally.

Phase 2 — Complete a single-cycle RV32I core first

  • Implement instruction decoding and the control unit.
  • Add operand, write-back, and next-PC selection logic.
  • Add byte enables plus load extraction/sign extension for LB/LH/LW/LBU/LHU and SB/SH/SW.
  • Implement LUI, AUIPC, JAL, and JALR.
  • Connect the existing blocks under an rv32i_core top level.
  • Verify all planned instructions before introducing pipeline hazards.

Exit criterion: directed programs execute correctly in a complete non- pipelined core.

Phase 3 — Introduce the five-stage pipeline

  • Add valid bits and IF/ID, ID/EX, EX/MEM, and MEM/WB registers.
  • Partition datapath and control signals by stage.
  • Carry the PC and PC + 4 to the stages that need them.
  • Confirm that retirement and memory side effects occur exactly once.

Exit criterion: hazard-free programs match the single-cycle reference.

Phase 4 — Resolve data and control hazards

  • Implement EX/MEM and MEM/WB forwarding for both ALU operands.
  • Forward store data and operands used by branch comparison where needed.
  • Detect load-use hazards and insert a one-cycle bubble.
  • Define branch/jump resolution stage and flush all younger wrong-path work.
  • Handle simultaneous stall, flush, and forwarding conditions explicitly.

Exit criterion: dependency-heavy and control-flow tests pass without manual NOP insertion.

Phase 5 — Architectural compliance and robustness

  • Add illegal-instruction handling or document the chosen behavior.
  • Decide and document misaligned-access behavior.
  • Run RV32I architectural compliance tests and differential tests.
  • Add SystemVerilog assertions and functional coverage for hazards and instruction classes.
  • Add continuous-integration lint and regression jobs.

Exit criterion: the defined RV32I subset passes its automated compliance and regression suite.

Phase 6 — Synthesis and optional extensions

  • Add a synthesis-friendly memory or external instruction/data bus.
  • Choose an FPGA board, add clock/reset wrappers and timing constraints, and check timing/resource reports.
  • Add memory-mapped UART/GPIO and a small software demo if desired.
  • Measure maximum frequency, CPI, and resource use.
  • Consider traps/interrupts, CSRs, Zicsr, M, or caches only after the baseline core is stable.

Exit criterion: a documented FPGA or ASIC-oriented build runs a repeatable demo and meets its stated timing target.

Known issues

  • Several parameterized modules currently terminate parameter declarations with semicolons instead of separating them with commas; this is invalid SystemVerilog syntax.
  • bramch_logic.sv is misspelled, although its module is named branch_logic.
  • The registered instruction and combinational pc_out in if_stage may refer to different fetch cycles. A proper IF/ID register should keep them together.
  • if_stage hard-codes 32-bit pc and imem instances even though the stage exposes a DATA_WIDTH parameter.
  • Reset signal names are inconsistent (rst_n and resN).
  • imem requires an absent working-directory-relative imem.hex file.
  • Memories have no reset/initial value policy, bounds checking, byte enables, or alignment checking.
  • dmem only supports word operations, so the full RV32I load/store set is not yet implemented.
  • The register file clears all entries in one asynchronous-reset loop. This is convenient for simulation but may not map efficiently to FPGA memory.
  • No block validates its raw control encoding, and there is no shared package to prevent encoding mismatches.
  • There is no license file yet; add one before defining reuse/distribution terms.

Suggested contribution workflow

Keep each change small and verifiable: add or update a test with every RTL change, run lint and the full regression, and document any new interface or control encoding. Completion should be judged by the exit criteria above, not only by the presence of modules.

About

A high-performance 5-stage pipelined RISC-V RV32I core implemented in SystemVerilog with full ISA compliance, hazard detection, and comprehensive verification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages