Team Name: __BLANK__
Course: CS209P - Computer Organization and Architecture
Institute: Indian Institute of Technology (IIT), Tirupati
Authors: Ch Pranav Tej (CS24B057) & Shivam Purve (CS24B055)
This project represents a complete, cycle-accurate software simulation of a 32-bit RISC-V (RV32I) processor. Developed over three phases, the simulator models the complex interaction between high-level pipeline scheduling and low-level memory physics.
- Phase 1 focused on the core computational engine: a 5-stage pipeline featuring hazard detection, branch flushing, and data forwarding.
- Phase 2 extended this engine into the memory domain, integrating a Harvard-architecture L1 cache and a unified L2 cache to model dynamic latency and structural backpressure.
- Phase 3 implemented a full Virtual Memory subsystem, introducing a Data TLB, a Hierarchical Two-Level Page Table, True LRU page eviction, and a trace-replay engine for simulating heavy OS-level memory workloads.
The resulting tool is a rigorous environment for observing how architectural decisions—like associativity, forwarding, or physical frame limits—mathematically impact the execution of real-world algorithms and massive memory traces.
To accurately simulate hardware parallelism using sequential C++ code, the simulator executes pipeline stages in reverse order: WriteBack → Memory → Execute → Decode → Fetch. This prevents an instruction from "rippling" through multiple stages in a single simulated cycle, ensuring that data only moves across stage latches (IF_ID, ID_EX, EX_MEM, MEM_WB) at the start of each new clock.
The HDU ensures architectural integrity by managing three primary hazard types:
- Load-Use Hazards: Detects when a load instruction is in the Execute stage and its destination register is required by the instruction currently in Decode, necessitating a 1-cycle stall.
- Control Hazards: Evaluates branches (
bne,blt,jal) in the Decode stage. If a branch is "Taken," the HDU flushes the instructions already fetched behind it. - Structural Hazards: Managed via countdown timers. If an instruction (like a multi-cycle
addi,mul, or a cache/VM miss) takes more than one cycle, the HDU freezes all preceding stages until the functional unit is free.
The simulator minimizes RAW (Read-After-Write) stalls by implementing a bypass network:
- EX-to-EX Forwarding: Passes results directly from the ALU output back to the ALU input.
- MEM-to-EX Forwarding: Passes data from the Memory stage back to the ALU input. (The Forwarding Unit can be toggled via the CLI to demonstrate the massive cycle-count delta in loop-heavy code).
The simulator implements a cycle-accurate model of the following hierarchy (PIPT - Physically Indexed, Physically Tagged):
- L1 Cache (Split):
L1I(Instruction) andL1D(Data) caches operate independently to allow simultaneous fetch and load/store operations. - L2 Cache (Unified): A larger shared cache that acts as the primary backup for L1 misses.
- Main Memory: The physical DRAM with a high cycle penalty.
Physical parameters are entirely driven by config.txt. The parser dynamically calculates set indices and tag bits based on cache sizes, block size, associativity, and latency values.
The Phase 3 extension introduces OS-level memory management, translating 32-bit Virtual Addresses to Physical Addresses before they hit the cache.
A fully-associative Data Translation Lookaside Buffer (DTLB) caches recent VPN-to-PFN translations.
- TLB Shootdown: If the Page Table evicts a physical frame to disk, the VM subsystem automatically queries the TLB and invalidates the corresponding entry to prevent stale translation corruption.
Instead of a memory-hogging flat array, the simulator utilizes a Two-Level Hierarchical Page Table (10-bit Directory Index, 10-bit Table Index). L2 page tables are dynamically allocated only when a page fault occurs, drastically reducing the memory footprint of the simulator.
Physical frames are managed via a strict Least Recently Used (LRU) policy.
- A global
access_counteracts as the master clock. - Even if the Page Table is bypassed by a TLB Hit, the TLB explicitly updates the physical frame's
last_usedtimestamp. - On eviction, the simulator checks the frame's
dirtybit and accurately increments the writeback counter if the page was modified.
.
├── config.txt # Dynamic geometry and latency configuration (INI-style)
├── main.cpp # Interactive CLI and Mode Selector
├── run_all.sh # Batch script to execute all Phase 3 trace files
├── simulator.cpp / .h # Core Pipeline, HDU, and Execution Engine
├── parser.cpp / .h # Parser for Assembly (.asm) and Trace (.trace) files
├── vm_config.h # VM parameter derivation
├── vm_subsystem.cpp / .h # VM Controller (TLB -> PT -> Penalty integration)
├── page_table.cpp / .h # Two-Level Page Table & Physical Frame Allocator
├── dtlb.cpp / .h # Data Translation Lookaside Buffer
├── cache_system.cpp / .h # L1/L2 Hierarchy Controller
├── cache.cpp / .h # Cache sets, logic, and replacement math
├── cache_block.cpp / .h # Cache line metadata structures
├── phase3_results/ # Output dumps of trace evaluation metrics
├── tests/ # Verified RISC-V assembly (.asm) test programs
├── MoM.md # Minutes of Meeting and internal docs
└── README.md # This documentation
Compile the simulator using C++17:
g++ -std=c++17 -o riscv_sim.exe *.cppRun the compiled binary:
./riscv_sim.exeUpon execution, the interactive CLI allows you to select between two modes:
- Executes
.asmfiles from thetests/directory (e.g.,tests/bubble_sort.asm). - Supports Data Forwarding toggles and Single-Step execution.
- In Single-Step mode, prints a visual layout of pipeline registers and Cache block states (Tags, Valid bits, LRU Age) every cycle.
- Ingests heavy memory traces (
L,S,ADD,MUL) using Virtual Addresses. - Automatically routes memory requests through the
VMSubsystembefore theCacheSystem. - Runs to completion and generates a detailed Execution Dashboard.
To run all 10 Phase 3 traces automatically and extract the metrics, execute the provided shell script:
chmod +x run_all.sh
./run_all.shAt the conclusion of a trace or assembly run, the simulator outputs an Execution Dashboard detailing the exact architectural penalties incurred:
- [PIPELINE STATISTICS]: Total Cycles, Instructions, Hazard Stalls, and overall IPC.
- [L1/L2 CACHE STATISTICS]: Total Accesses, Hits, Misses, and Miss Rates for the memory hierarchy.
- [VIRTUAL MEMORY STATISTICS]: Highlights TLB Hits/Misses, Page Faults, Frame Evictions, Dirty Writebacks, and the total Translation Penalty cycles injected into the pipeline memory stage. (Traces 06, 08, and 09 explicitly demonstrate memory "thrashing" when the virtual working set vastly exceeds the physical frame limit).