Which LLM understands the math, and which is faking it? A competition on robust vs. spurious reasoning in language models. @neuripsconf.bsky.social 2026 ๐ aimo-interp.github.io
Nothing here yet.
Which LLM understands the math, and which is faking it? A competition on robust vs. spurious reasoning in language models. @neuripsconf.bsky.social 2026 ๐ aimo-interp.github.io
Nothing here yet.