Skip to content

Repository files navigation

CS209P Project: Integrated RISC-V 5-Stage Pipelined Simulator with 2-Level Cache & Virtual Memory

Team Name: __BLANK__
Course: CS209P - Computer Organization and Architecture
Institute: Indian Institute of Technology (IIT), Tirupati
Authors: Ch Pranav Tej (CS24B057) & Shivam Purve (CS24B055)


1. Project Overview

This project represents a complete, cycle-accurate software simulation of a 32-bit RISC-V (RV32I) processor. Developed over three phases, the simulator models the complex interaction between high-level pipeline scheduling and low-level memory physics.

  • Phase 1 focused on the core computational engine: a 5-stage pipeline featuring hazard detection, branch flushing, and data forwarding.
  • Phase 2 extended this engine into the memory domain, integrating a Harvard-architecture L1 cache and a unified L2 cache to model dynamic latency and structural backpressure.
  • Phase 3 implemented a full Virtual Memory subsystem, introducing a Data TLB, a Hierarchical Two-Level Page Table, True LRU page eviction, and a trace-replay engine for simulating heavy OS-level memory workloads.

The resulting tool is a rigorous environment for observing how architectural decisions—like associativity, forwarding, or physical frame limits—mathematically impact the execution of real-world algorithms and massive memory traces.


2. Pipeline Architecture (Phase 1 Foundations)

2.1 Reverse-Order Execution Logic

To accurately simulate hardware parallelism using sequential C++ code, the simulator executes pipeline stages in reverse order: WriteBack → Memory → Execute → Decode → Fetch. This prevents an instruction from "rippling" through multiple stages in a single simulated cycle, ensuring that data only moves across stage latches (IF_ID, ID_EX, EX_MEM, MEM_WB) at the start of each new clock.

2.2 Hazard Detection Unit (HDU)

The HDU ensures architectural integrity by managing three primary hazard types:

  • Load-Use Hazards: Detects when a load instruction is in the Execute stage and its destination register is required by the instruction currently in Decode, necessitating a 1-cycle stall.
  • Control Hazards: Evaluates branches (bne, blt, jal) in the Decode stage. If a branch is "Taken," the HDU flushes the instructions already fetched behind it.
  • Structural Hazards: Managed via countdown timers. If an instruction (like a multi-cycle addi, mul, or a cache/VM miss) takes more than one cycle, the HDU freezes all preceding stages until the functional unit is free.

2.3 Data Forwarding Unit

The simulator minimizes RAW (Read-After-Write) stalls by implementing a bypass network:

  • EX-to-EX Forwarding: Passes results directly from the ALU output back to the ALU input.
  • MEM-to-EX Forwarding: Passes data from the Memory stage back to the ALU input. (The Forwarding Unit can be toggled via the CLI to demonstrate the massive cycle-count delta in loop-heavy code).

3. Memory Hierarchy & Cache Design (Phase 2)

3.1 Three-Tier Hierarchy

The simulator implements a cycle-accurate model of the following hierarchy (PIPT - Physically Indexed, Physically Tagged):

  • L1 Cache (Split): L1I (Instruction) and L1D (Data) caches operate independently to allow simultaneous fetch and load/store operations.
  • L2 Cache (Unified): A larger shared cache that acts as the primary backup for L1 misses.
  • Main Memory: The physical DRAM with a high cycle penalty.

3.2 Dynamic Hardware Configuration

Physical parameters are entirely driven by config.txt. The parser dynamically calculates set indices and tag bits based on cache sizes, block size, associativity, and latency values.


4. Virtual Memory Subsystem (Phase 3)

The Phase 3 extension introduces OS-level memory management, translating 32-bit Virtual Addresses to Physical Addresses before they hit the cache.

4.1 Data TLB & Shootdown Coherence

A fully-associative Data Translation Lookaside Buffer (DTLB) caches recent VPN-to-PFN translations.

  • TLB Shootdown: If the Page Table evicts a physical frame to disk, the VM subsystem automatically queries the TLB and invalidates the corresponding entry to prevent stale translation corruption.

4.2 Dynamic Two-Level Page Table

Instead of a memory-hogging flat array, the simulator utilizes a Two-Level Hierarchical Page Table (10-bit Directory Index, 10-bit Table Index). L2 page tables are dynamically allocated only when a page fault occurs, drastically reducing the memory footprint of the simulator.

4.3 True LRU Eviction & Dirty Writebacks

Physical frames are managed via a strict Least Recently Used (LRU) policy.

  • A global access_counter acts as the master clock.
  • Even if the Page Table is bypassed by a TLB Hit, the TLB explicitly updates the physical frame's last_used timestamp.
  • On eviction, the simulator checks the frame's dirty bit and accurately increments the writeback counter if the page was modified.

5. Repository Structure

.
├── config.txt                 # Dynamic geometry and latency configuration (INI-style)
├── main.cpp                   # Interactive CLI and Mode Selector
├── run_all.sh                 # Batch script to execute all Phase 3 trace files
├── simulator.cpp / .h         # Core Pipeline, HDU, and Execution Engine
├── parser.cpp / .h            # Parser for Assembly (.asm) and Trace (.trace) files
├── vm_config.h                # VM parameter derivation
├── vm_subsystem.cpp / .h      # VM Controller (TLB -> PT -> Penalty integration)
├── page_table.cpp / .h        # Two-Level Page Table & Physical Frame Allocator
├── dtlb.cpp / .h              # Data Translation Lookaside Buffer
├── cache_system.cpp / .h      # L1/L2 Hierarchy Controller
├── cache.cpp / .h             # Cache sets, logic, and replacement math
├── cache_block.cpp / .h       # Cache line metadata structures
├── phase3_results/            # Output dumps of trace evaluation metrics
├── tests/                     # Verified RISC-V assembly (.asm) test programs
├── MoM.md                     # Minutes of Meeting and internal docs
└── README.md                  # This documentation

6. Build and Execution Instructions

6.1 Build Command

Compile the simulator using C++17:

g++ -std=c++17 -o riscv_sim.exe *.cpp

6.2 Execution Modes

Run the compiled binary:

./riscv_sim.exe

Upon execution, the interactive CLI allows you to select between two modes:

[1] Assembly Mode (Phases 1 & 2):

  • Executes .asm files from the tests/ directory (e.g., tests/bubble_sort.asm).
  • Supports Data Forwarding toggles and Single-Step execution.
  • In Single-Step mode, prints a visual layout of pipeline registers and Cache block states (Tags, Valid bits, LRU Age) every cycle.

[2] Trace Replay Mode (Phase 3):

  • Ingests heavy memory traces (L, S, ADD, MUL) using Virtual Addresses.
  • Automatically routes memory requests through the VMSubsystem before the CacheSystem.
  • Runs to completion and generates a detailed Execution Dashboard.

Batch Trace Evaluation

To run all 10 Phase 3 traces automatically and extract the metrics, execute the provided shell script:

chmod +x run_all.sh
./run_all.sh

7. Performance Metrics & Analytics Dashboard

At the conclusion of a trace or assembly run, the simulator outputs an Execution Dashboard detailing the exact architectural penalties incurred:

  • [PIPELINE STATISTICS]: Total Cycles, Instructions, Hazard Stalls, and overall IPC.
  • [L1/L2 CACHE STATISTICS]: Total Accesses, Hits, Misses, and Miss Rates for the memory hierarchy.
  • [VIRTUAL MEMORY STATISTICS]: Highlights TLB Hits/Misses, Page Faults, Frame Evictions, Dirty Writebacks, and the total Translation Penalty cycles injected into the pipeline memory stage. (Traces 06, 08, and 09 explicitly demonstrate memory "thrashing" when the virtual working set vastly exceeds the physical frame limit).

About

A cycle-accurate 32-bit RISC-V (RV32I) simulator in C++ featuring a 5-stage pipeline, dynamically configurable 2-level cache hierarchy, and a full Virtual Memory subsystem with a Data TLB and Two-Level Page Table

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages