Currently excited about studying the AI alignment problem and exploring LLM interpretability, evaluations and oversight.
- Gaming the Answer Matcher
- Lock In Risk Benchmark
- Semantic Communication
- Legal Agentic Evals
- Scalable Oversight, LLM Evaluators
- MITR Take Home Eval
| Of Epics, Networks and Bots (May 19, 2024) | |
| A Statistical Dive into The Ramayan: Code Walkthrough Part 1 (Jun 5, 2024) | |
| A Statistical Dive into The Ramayan: Code Walkthrough Part 2 (Jun 5, 2024) |
