NEWESTPODCAST
Can agents sabotage code and safety research?
What unprompted evaluations, continuation tests, refusal patterns, and model-organism audits actually establish
Show notes & transcriptBlog
Podcasts
- Can agents sabotage code and safety research?
- Overt failure, covert failure, and stealth
- Anatomy of the Anthropic covert-sabotage case
- Covert Sabotage — Topic Overview
Projects
- fidx
local hybrid semantic search that beats QMD on recall and precision at ~300–1000× lower latency.
- token-count-compare
a side-by-side tokenizer comparison across frontier models.
- token-usage-analyzer
finds what's burning your token budget in AI coding CLIs.