We've been monitoring a decline in HuggingFace downloads across the top open LLMs we track for interconnects. It looks like there was a slight change in counting of downloads, as we saw a ~30% download rate drop for many top models, but this is continuing over time. dashboard.interconnects.ai
OSS Safety Models
by @julietshen.online
a feed to learn, discuss, and rant about open source safety models
De-skilling myself on doing CAPTCHAS and navigating dark patterns in customer service.
the latest anthropic cry abt "cruelty" & "abuse" of ai is truly stupid. i beg journalists to stop parroting amodei we wrote why dehumanisation of ai is impossible & 6+ yrs later, everything in this paper holds Robot Rights? Let’s Talk about Human Welfare Instead dl.acm.org/doi/pdf/10.1...
I’m more bought into AI becoming far more efficient, which is a massive lever on diffusion, than I am on AI models making massive breakthroughs that result in eliminating all the known weaknesses of LLMs. www.interconnects.ai/p/i-expect-r...
Randomized trials with GPT-4o: "AI access raises test scores, and a smaller gain persists a week later. Gains remain for students who use AI as a tutor (“augmentation”) and fade for students who have AI write for them (“automation”)" Across all experiments, positive effects arxiv.org/abs/2607.08849
i appreciate the authors being upfront about the limitation of their work but what can be meaningfully inferred from work like this
We’re sharing the engineering work behind Ai2’s GPU scheduler—our new approach to allocating compute across research teams. On our largest H100 cluster, median queue wait fell from 5 minutes to 24 seconds. 🧵 buff.ly/4UsJEYe
From The Lancet: in an urgent care setting, the advice of the obsolete Gemini 2.5 Pro & Gemini 2.5 Flash (without access to patient medical records) were rated of similar quality to doctors by other physicians. There was no safety issues spotted. Models have gotten significantly better since.
This document is going to be an assigned reading in college classes that cover this moment in time, there's a lot happening in a few paragraphs... www.ahmath.org/statements
we have a policy for the use of genAI at the @aial.ie. we reflect on values key to scientific integrity, such as replicability, and err on the side of not using it blogpost: aial.ie/blog/gen-ai-... pdf: zenodo.org/records/2322...
As a researcher who did early some work on the productivity impacts of AI chatbots using RCTs, I’d note a lack of similar studies since the dawn of true agents last fall Partially that is newness & partially research design challenges, but I suspect we are missing some large & important effects
Some early first-hand accounts of the experience of encountering a narrow superhuman intelligence as mathematicians grapple with the hundreds of big AI proofs released by OpenAI. Problems solved in inhuman ways that make us wonder what it means to actually know things... scottaaronson.blog?p=10169