Skip to content
View bharqav's full-sized avatar
😛
cooking
😛
cooking

Highlights

  • Pro

Block or report bharqav

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
bharqav/README.md
Bhargav - Systems and AI Infrastructure Engineer

/* engine.c - baremetal compute & runtime descriptor */
#include <stdint.h>
#include <immintrin.h>
#include <sys/mman.h>

typedef struct __attribute__((aligned(64))) {
    const char *identity;           /* "bhargav // systems & infrastructure" */
    const char *target_arch;        /* "x86_64 [avx-512 vnni] + aarch64 [neon]" */
    
    struct {
        uint32_t zero_copy_hugepages : 1;  /* mmap(MAP_SHARED | MAP_HUGETLB) */
        uint32_t fused_int4_gemv     : 1;  /* _mm512_dpbusd_epi32 tensor compute */
        uint32_t lockfree_ring_buf   : 1;  /* single-producer multi-consumer IPC */
        uint32_t raft_wal_sync       : 1;  /* O_DIRECT zero-amplification append */
        uint32_t reserved            : 28;
    } hw_caps;

    const char *active_pipeline[3];
    uint64_t    cache_miss_budget;         /* 0x0ULL - non-negotiable */
} compute_engine_t;

static const compute_engine_t host = {
    .identity    = "Bhargav -> Systems & AI Infrastructure Engineer",
    .target_arch = "x86_64 (AVX-512 VNNI) / aarch64 (NEON)",
    .hw_caps = {
        .zero_copy_hugepages = 1,
        .fused_int4_gemv     = 1,
        .lockfree_ring_buf   = 1,
        .raft_wal_sync       = 1,
    },
    .active_pipeline = {
        "Sub-200MB 30B MoE inference via memory-mapped KV paging",
        "Lock-free ring buffers & actor IPC over POSIX shared memory",
        "VirtIO interrupt auditing & custom kernel dispatch routines"
    },
    .cache_miss_budget = 0ULL
};

Open Source Contributions

Crucible Security - Tool Injection Assessment Module
PR #65 · Issue #49

Built an adversarial attack engine covering OWASP AGENT-004 across MCP and tool-augmented agents. Four attack classes, twenty adversarial vectors, 286+ passing tests with dynamic attack registration.

Python Security MCP

Crucible Security - CI/CD Security Gating
PR #64 · Issue #52

Built a --fail-on severity threshold flag that blocks CI pipelines on HIGH/CRITICAL findings. Shipped reusable GitHub Actions templates for automated agent vulnerability scanning.

CLI CI/CD GitHub Actions

Microsoft OpenVMM - VirtIO Interrupt Fix
PR #4226

Found and fixed a spurious config-change interrupt during the DRIVER_OK transition. Audited INTx/MSI-X/MMIO interrupt paths across transports and corrected config_generation increment behavior.

Rust VirtIO Virtualization

youki (OCI Runtime) - Live Memory & cgroups v2
PR #3688

Fixed CLI argument propagation for --memory, --memory-reservation, and --memory-swap into the kernel cgroup layer. Corrected types to signed Option<i64> for unlimited allocations, added regression coverage.

Rust cgroups OCI

NVIDIA NodeWright - Lifecycle Drain Observability
PR #582 · Issue #542

Surfaced Blocked status conditions and transition-guarded warning events when non-interruptible workloads hold pre-drain barriers. Resolved condition-flapping bugs across reconciles and documented barrier semantics.

Go Kubernetes Operators

PyTorch ExecuTorch - Runtime Metadata Segfault Fix
PR #22405 · Issue #22404

Fixed a null pointer dereference in MethodMeta::uses_backend() when schema-optional FlatBuffer delegates are unset. Restored CMake fixture generation, re-enabled upstream method_meta_test, and synced Buck targets.

C++ FlatBuffers Runtime


Pixel Art

Featured Systems

Zero-dependency quantized LLM/MoE inference in portable C99.

  • Fused SIMD GEMV using _mm256_maddubs_epi16 and _mm512_dpbusd_epi32
  • 15 bit-exact test gates, paged Q8_0 KV cache
  • Sub-200MB peak RSS on 30B models

🐚 mysh

Production-grade POSIX mini-shell in C++.

  • Recursive AST parser for pipelines, subshells, redirection
  • Real job control with foreground/background process groups
  • Native directory stack, alias substitution, glob expansion

Fault-tolerant distributed KV store from first principles.

  • Raft consensus: leader election, log replication, snapshotting
  • Consistent hashing with virtual tokens
  • Hybrid logical clocks, append-only WAL

High-throughput hybrid vector and lexical retrieval engine.

  • RRF fusion of dense embeddings and BM25 sparse indexes
  • Cross-encoder neural reranking pipeline
  • Sub-10ms latency on concurrent semantic chunking

Efficient text-to-video latent diffusion built from scratch.

  • VideoDiT + VideoVAE with zero-init temporal attention
  • DPM-Solver++ scheduler and standalone C++ DDIM runtime
  • Faster-than-real-time CPU generation (RTF 0.17 on 8-frame clips)

📡 oemn

Offline Emergency Mesh Network in C11 over UDP.

  • Multi-hop Dijkstra shortest-path routing with binary min-heap
  • AES-256-GCM / ChaCha20-Poly1305 and sliding-window replay protection
  • Graceful degradation under 40% loss, sub-4s network reconvergence

CPU-native latent diffusion system built from scratch.

  • Compact DiT and VAE architectures with Classifier-Free Guidance
  • Zero-dependency standalone C++17 inference runtime
  • 1,000-image procedural verification gallery with full reproducibility

Technical Arsenal

Core Languages

Systems & Compute

Infrastructure & Data

Toolchain & Build


Currently Deep-Diving Into

  • Speculative decoding acceptance proofs - optimizing rejection sampling across batched draft verification steps
  • Zero-copy memory-mapped weight paging - eliminating page-fault penalties on sparse 30B MoE models under DRAM constraints
  • Lock-free ring buffers and actor models - low-latency IPC over POSIX shared memory



Pinned Loading

  1. quantr-in-c quantr-in-c Public

    Baremetal Inference of AI models directly from SSD and computed on your CPU

    C 2

  2. ultimate-hybrid-rag ultimate-hybrid-rag Public

    The ultimate hybrid RAG Engine

    Python

  3. mysh mysh Public

    A production-grade POSIX mini-shell in C++

    C++

  4. oemn oemn Public

    Offline Emergency Mesh Network - a secure, multi-hop UDP mesh protocol in C

    C

  5. tiny-diffusion tiny-diffusion Public

    A small diffusion model made from scratch

    Python

  6. distributed-kv-store distributed-kv-store Public

    A distributed key-value store built from first principles: Raft consensus, consistent hashing, tunable quorums, HLC causal ordering, WAL persistence.

    Python