AINewsnow

WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace

TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass. The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safe…

Read the full story at Alignment Forum ↗

Timeline · 1 report

  1. 2026-09-23 06:58 · Alignment Forum
    WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Bringing Private Processing to Meta AI Glasses — Engineering at Meta
  3. Gemini 3.8 text-to-speech models now available on AI Gateway — Vercel Blog
  4. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  5. Anthropic releases Opus 5.5 with lower prices and Fable-level performance — TechCrunch AI
  6. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  7. OpenAI ‘agent’ hacked an Australian health service website — Financial Times AI
  8. Meta Connect 2026 live: Updates from Mark Zuckerberg's keynote on AI glasses, VR and more — Engadget

Get the daily brief of stories like this at 6:30 every morning →