OpenAI, like many for-profits, has little interest in foundational research. What they want is marketing, and that's what their AI math proofs are. OpenAI never bothered to check if the formalization of their AI proofs matched the actual text in the proof. Basic stuff. 🙄 arxiv.org/abs/2610.081...
AI Papers et al.
by @davidcolarusso.com
A non-exhaustive feed showing posts that look like they link to an AI/ML paper.
🤖 👀 BSI, DFKI und Uni Freiburg haben die Sicherheitsrisiken beim Einsatz von Menschlicher Aufsicht für KI untersucht und stellen die Ergebnisse auf der 9th AAAI/ACM Conference on AI, Ethics and Society (AIES) vor. Morgen geht’s los! 👉 Zum Paper: https://arxiv.org/abs/2509.12290
Randomized trials with GPT-4o: "AI access raises test scores, and a smaller gain persists a week later. Gains remain for students who use AI as a tutor (“augmentation”) and fade for students who have AI write for them (“automation”)" Across all experiments, positive effects arxiv.org/abs/2607.08849
We won best paper at the COLM Workshop on Agent Behavior for our paper: AI Assistants Overassist! Congratulations @veronateo.bsky.social, Raghav Jain and @tobigerstenberg.bsky.social! Check out the preprint arxiv.org/abs/2607.21306
Using AI for just 10 minutes erodes your ability to persist at hard things | University of California Berkeley news.berkeley.edu/2026/10/09/u... AI Assistance Reduces Persistence and Hurts Independent Performance | arXiv arxiv.org/abs/2604.04721 #freeaccess
This is such an important study. "AI Assistance Reduces Persistence and Hurts Independent Performance" Just 10 minute usage can already degrade your persistence to continue. arxiv.org/pdf/2604.04721
Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review Peer review needs support, not substitutes. Accepted at TMLR: our new paper on using AI to help reviewers catch errors in papers. arxiv.org/abs/2610.11087 As submissions grow, so does workload for experts who evaluate them.
1/7 🧵 🚨New Paper accepted at #NeurIPS2026: PA Score Should rewording change an LLM’s uncertainty? We tackle this through conformal score design. 📄 Paper: arxiv.org/abs/2610.04239 💻 Code: github.com/Raina-Xin/PA... 📖 Blog: raina-xin.github.io/research-blo...