June 21, 2025Research Report3 min read
zebra_simple: Zebra Puzzle Test for LLMs
The zebra_simple benchmark compares models on the same Zebra Puzzle tasks (pure logic reasoning) and extracts their chains of thought for further analysis.
Read it →Notes
What we've written down as we work: benchmarks we publish, and critical reads of other people's takes on AI.
June 21, 2025Research Report3 min read
The zebra_simple benchmark compares models on the same Zebra Puzzle tasks (pure logic reasoning) and extracts their chains of thought for further analysis.
Read it →June 20, 2025Blog Post5 min read
A critical look at the definitions and conclusions in Yuval Noah Harari's video about AI and human trust.
Read it →