"Training Verifiably Robust Agnets Using Set-Based Reinforcement Learning", Wendl et. al, TMLR, 2026.
-
Updated
Jun 30, 2026 - Python
"Training Verifiably Robust Agnets Using Set-Based Reinforcement Learning", Wendl et. al, TMLR, 2026.
Code and resources for our paper on principled design for trustworthy AI: interpretability, robustness, and safety across modalities. ICLR 2026 Trustworthy AI Workshop, Rio de Janeiro.
Creates the ideal folder structure based on Principled AI
Semantic-layer prompt injection defence that separates untrusted instructions from authority while preserving the legitimate task.
"Training Verifiably Robust Agnets Using Set-Based Reinforcement Learning", Wendl et. al, TMLR, 2026.
Add a description, image, and links to the principled-ai topic page so that developers can more easily learn about it.
To associate your repository with the principled-ai topic, visit your repo's landing page and select "manage topics."