GenAI Engineer Intern @ AllCognix AI, building RAG pipelines and production ML infrastructure on Haystack 2.x.
- 15 merged PRs across deepset-ai/haystack (25k★), run-llama/llama_index (40k★), mlflow, pydata/sparse and kubeflow/mcp-server, plus 10 more in review across Haystack, Kubeflow, Jaeger and AiiDA
- Built ShardFlow, distributed LLM inference hitting 28.10 TPS on Qwen2.5-7B across free Kaggle GPUs over public WAN
- Cut RAG latency by 40% and token costs by 60% via Haystack 2.x migration + context windowing
- Reduced document ingestion from 70s → 27s with parallel processing
- Published quantization stability research on Jetson Nano edge hardware: preprint on Zenodo
|
High-performance distributed LLM inference framework. Partitions any HuggingFace transformer across N heterogeneous GPU machines over public internet WAN using neural speculative decoding, zero-copy binary tensor serialization, and a Rust TCP relay on AWS EC2.
28.10 TPS peak · 5.71x speedup · K=8 speculative decoding · 65% draft accept rate · Qwen2.5-7B across free Kaggle T4s |
Observability and diagnostics engine for Haystack 2.x RAG pipelines. Exposes document-store validation, pipeline inspection, retrieval-failure analysis, and structured debug bundle diffing via MCP.
6-class retrieval-failure taxonomy · 823-chunk live corpus · ~0.95s for 15 concurrent MCP requests |
|
Production ML platform with ConvNeXt-Tiny inference, CLIP-based input validation, OpenCV symptom overlays, GPT-4o mini recommendations, and a human-in-the-loop feedback pipeline.
88.46% accuracy · 75% model compression · 65ms CPU inference |
Time-series sales forecasting on 1M+ rows of Rossmann store data. Feature engineering on promotions, holidays, and store metadata. Containerized and served via FastAPI.
1M+ rows · containerized API · promotion and seasonality features |
15 merged PRs across deepset-ai/haystack (25k★), run-llama/llama_index (40k★), mlflow/mlflow, pydata/sparse and kubeflow/mcp-server, with 10 more open across Haystack, Kubeflow, Jaeger and AiiDA.
| PR | Repo | Change |
|---|---|---|
| #12529 | haystack | DocumentSplitter: new split_by="token" mode using tiktoken |
| #12485 | haystack | ComponentTool deserialization support in OpenAI and Azure responses chat generators |
| #12407 | haystack | AzureOpenAIChatGenerator.to_dict(): TypeError when response_format is a plain dict |
| #12387 | haystack | PipelineBase.__eq__: unhandled AssertionError on non-Pipeline comparison |
| #12206 | haystack | PipelineBase.remove_component: auto-variadic socket state not restored on removal |
| #11987 | haystack | EmbeddingBasedDocumentSplitter: split_idx_start metadata not populated |
| #11847 | haystack | FallbackChatGenerator: nested chat generators lost on to_dict() roundtrip |
| #11768 | haystack | RecursiveDocumentSplitter: split_overlap ignored on no-separator fallback path |
| #11711 | haystack | RecursiveDocumentSplitter: wrong split_idx_start with word/token units and overlap |
| #11419 | haystack | DocumentLanguageClassifier: crash on docs with content=None |
| #22167 | llama_index | SemanticDoubleMergingSplitterNodeParser: stopword removal now uses word tokenization |
| #25862 | mlflow | set_logged_model_tags: bulk upsert for SQLite, PostgreSQL and MySQL |
| #274 | kubeflow/mcp-server | Unrestricted access not propagated when inheriting from parent persona |
| #248 | kubeflow/mcp-server | Return RESOURCE_NOT_FOUND on missing job in trainer monitoring tools |
| #279 | kubeflow/mcp-server | Make _inject_trainer_hf_home thread-safe for concurrent calls |
| #955 | pydata/sparse | sparse.diagonal: support for negative offsets, rectangular shapes and negative axes |
| #960 | pydata/sparse | save_npz / load_npz: support for CSR, CSC and DOK formats |
| PR | Repo | Change |
|---|---|---|
| #12990 | haystack | Agent: schema-constrained structured outputs with a recovery loop |
| #12824 | haystack | Redact ImageContent and FileContent in ToolCallResult trace dicts |
| #806 | kubeflow/sdk | get_container_devices: handle empty and memory-only resource limits |
| #3175 | kubeflow/spark-operator | ScheduledSparkApplication: recover from FailedValidation once spec is fixed |
| #199 | kubeflow/pipelines-components | Handle empty metadata.yaml in check_component_freshness |
| #9658 | jaeger | ai-sidecar: point default MCP URL to query port 16686 |
| #9511 | jaeger | ai-sidecar: join ACP prompt blocks with a delimiter to prevent token fusion |
| #7598 | aiida-core | JsonableData: preserve @module and @class keys before from_dict() |
- Interning @ AllCognix AI: production RAG system on Haystack 2.x, EC2, Weaviate, Redis, Vault
- OSS contributions to deepset-ai/haystack (10 merged PRs, more in review) and Kubeflow (mcp-server, sdk, spark-operator, pipelines-components)
- Research paper targeting Computers and Electronics in Agriculture (Q1 Elsevier)
- Open to ML Engineer / GenAI roles: remote, India & international


