Skip to content
View rautaditya2606's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report rautaditya2606

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rautaditya2606/README.md

LinkedIn Portfolio Email


Building production ML systems that actually ship — from edge devices to distributed GPU clusters.


About

GenAI Engineer Intern @ AllCognix AI, building RAG pipelines and production ML infrastructure on Haystack 2.x.

  • 15 merged PRs across deepset-ai/haystack (25k★), run-llama/llama_index (40k★), mlflow, pydata/sparse and kubeflow/mcp-server, plus 10 more in review across Haystack, Kubeflow, Jaeger and AiiDA
  • Built ShardFlow, distributed LLM inference hitting 28.10 TPS on Qwen2.5-7B across free Kaggle GPUs over public WAN
  • Cut RAG latency by 40% and token costs by 60% via Haystack 2.x migration + context windowing
  • Reduced document ingestion from 70s → 27s with parallel processing
  • Published quantization stability research on Jetson Nano edge hardware: preprint on Zenodo

Tech Stack

ML / Deep Learning

PyTorch Scikit-learn LightGBM ONNX

MLOps & Deployment

FastAPI Docker GitHub Actions AWS

Backend & Data

Python PostgreSQL Rust


Projects

High-performance distributed LLM inference framework. Partitions any HuggingFace transformer across N heterogeneous GPU machines over public internet WAN using neural speculative decoding, zero-copy binary tensor serialization, and a Rust TCP relay on AWS EC2.

PyTorch HuggingFace CUDA Graphs AWS EC2 Rust

28.10 TPS peak · 5.71x speedup · K=8 speculative decoding · 65% draft accept rate · Qwen2.5-7B across free Kaggle T4s

Observability and diagnostics engine for Haystack 2.x RAG pipelines. Exposes document-store validation, pipeline inspection, retrieval-failure analysis, and structured debug bundle diffing via MCP.

Haystack 2.x Weaviate MCP Python

6-class retrieval-failure taxonomy · 823-chunk live corpus · ~0.95s for 15 concurrent MCP requests

Production ML platform with ConvNeXt-Tiny inference, CLIP-based input validation, OpenCV symptom overlays, GPT-4o mini recommendations, and a human-in-the-loop feedback pipeline.

FastAPI ONNX Runtime PostgreSQL Docker OpenAI

88.46% accuracy · 75% model compression · 65ms CPU inference

Time-series sales forecasting on 1M+ rows of Rossmann store data. Feature engineering on promotions, holidays, and store metadata. Containerized and served via FastAPI.

LightGBM FastAPI Docker Pandas

1M+ rows · containerized API · promotion and seasonality features


Open Source

Open Source

15 merged PRs across deepset-ai/haystack (25k★), run-llama/llama_index (40k★), mlflow/mlflow, pydata/sparse and kubeflow/mcp-server, with 10 more open across Haystack, Kubeflow, Jaeger and AiiDA.

Merged

PR Repo Change
#12529 haystack DocumentSplitter: new split_by="token" mode using tiktoken
#12485 haystack ComponentTool deserialization support in OpenAI and Azure responses chat generators
#12407 haystack AzureOpenAIChatGenerator.to_dict(): TypeError when response_format is a plain dict
#12387 haystack PipelineBase.__eq__: unhandled AssertionError on non-Pipeline comparison
#12206 haystack PipelineBase.remove_component: auto-variadic socket state not restored on removal
#11987 haystack EmbeddingBasedDocumentSplitter: split_idx_start metadata not populated
#11847 haystack FallbackChatGenerator: nested chat generators lost on to_dict() roundtrip
#11768 haystack RecursiveDocumentSplitter: split_overlap ignored on no-separator fallback path
#11711 haystack RecursiveDocumentSplitter: wrong split_idx_start with word/token units and overlap
#11419 haystack DocumentLanguageClassifier: crash on docs with content=None
#22167 llama_index SemanticDoubleMergingSplitterNodeParser: stopword removal now uses word tokenization
#25862 mlflow set_logged_model_tags: bulk upsert for SQLite, PostgreSQL and MySQL
#274 kubeflow/mcp-server Unrestricted access not propagated when inheriting from parent persona
#248 kubeflow/mcp-server Return RESOURCE_NOT_FOUND on missing job in trainer monitoring tools
#279 kubeflow/mcp-server Make _inject_trainer_hf_home thread-safe for concurrent calls
#955 pydata/sparse sparse.diagonal: support for negative offsets, rectangular shapes and negative axes
#960 pydata/sparse save_npz / load_npz: support for CSR, CSC and DOK formats

In Review

PR Repo Change
#12990 haystack Agent: schema-constrained structured outputs with a recovery loop
#12824 haystack Redact ImageContent and FileContent in ToolCallResult trace dicts
#806 kubeflow/sdk get_container_devices: handle empty and memory-only resource limits
#3175 kubeflow/spark-operator ScheduledSparkApplication: recover from FailedValidation once spec is fixed
#199 kubeflow/pipelines-components Handle empty metadata.yaml in check_component_freshness
#9658 jaeger ai-sidecar: point default MCP URL to query port 16686
#9511 jaeger ai-sidecar: join ACP prompt blocks with a delimiter to prevent token fusion
#7598 aiida-core JsonableData: preserve @module and @class keys before from_dict()

Currently

  • Interning @ AllCognix AI: production RAG system on Haystack 2.x, EC2, Weaviate, Redis, Vault
  • OSS contributions to deepset-ai/haystack (10 merged PRs, more in review) and Kubeflow (mcp-server, sdk, spark-operator, pipelines-components)
  • Research paper targeting Computers and Electronics in Agriculture (Q1 Elsevier)
  • Open to ML Engineer / GenAI roles: remote, India & international


Pune, India · B.Tech CSE (AI & Analytics) · MIT ADT University · 2028

Pinned Loading

  1. Shardflow Shardflow Public

    Python 3

  2. haystack-diagnostics haystack-diagnostics Public

    Python 6

  3. research_paper research_paper Public

    Jupyter Notebook

  4. wheat_detection wheat_detection Public

    Python 2

  5. Rossman-Deployed Rossman-Deployed Public

    Jupyter Notebook 2

  6. FastAPI_NYC FastAPI_NYC Public

    Jupyter Notebook