Add Practical Lessons for GenAI Evals with Chip Huyen and Vivienne Zhang - #97
Open
conorbronsdon wants to merge 2 commits into
Open
conorbronsdon wants to merge 2 commits into
conorbronsdon wants to merge 2 commits into
Conversation
conorbronsdon
marked this pull request as ready for review
September 10, 2026 00:34
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add a Chain of Thought Podcast panel with Chip Huyen and Vivienne Zhang to the podcast section. The episode gives practitioners concrete ways to evaluate agent plans and tool use.
Why it belongs
This is an eval-first episode with a public transcript. At 28:00, Huyen explains why generic labels such as coherence and relevance are not stable metrics without clear criteria, fixed judge prompts, and curated examples of good and bad responses. At 39:00, she breaks agent evaluation into task completion under constraints, valid tool selection and parameters, plan efficiency, and metrics for the application's likely failure points. Zhang then discusses human handoffs at 42:00 and applying evals at the failure points in a customer-service workflow.
Disclosure
I host Chain of Thought and am submitting my own episode. At the time of recording, I was Head of Developer Awareness at Galileo. Galileo sponsored the episode; its promotional read begins at 20:00. The panel was recorded at Galileo's Productionize 2.0 event and was moderated by Galileo co-founder and CTO Atindriyo Sanyal.
Verification
README.md.