Can @canrager.bsky.social
@tamarott.bsky.social @tamarott.bsky.social
Arnab Sen Sharma @arnabsensharma.bsky.social PhD Student at Northeastern, working to make LLMs interpretable
BlackboxNLP @blackboxnlp.bsky.social The largest workshop on analysing and interpreting neural networks for NLP.
BlackboxNLP will be held at EMNLP 2025 in Suzhou, China
blackboxnlp.github.io
Hiba Ahsan @hibaahsan.bsky.social PhD student @ Northeastern University, Clinical NLP
https://hibaahsan.github.io/
she/her
Hye Sun Yun @hyesunyun.bsky.social PhD candidate in CS at Northeastern University | NLP + HCI for health | she/her 🏃♀️🧅🌈
Yida Chen @yidachen.bsky.social CS PhD student at Harvard. Interested in Interpretability 🔍, Visualizations 📊, Human-AI Interaction🧍🤖. All opinions are mine. https://yc015.github.io/
Chantal @chantalsh.bsky.social PhD (in progress) @ Northeastern! NLP 🤝 LLMs
she/her
Julius Adebayo @juliusad.bsky.social ML researcher, building interpretable models at Guide Labs (guidelabs.bsky.social).
Maxime Méloux @maximemeloux.bsky.social PhD student @LIG | Causal abstraction, interpretability & LLMs
Calvin McCarter @calvinmccarter.bsky.social calvinmccarter.com
Derek Parfait @cole-haus.bsky.social Trying to figure things out about how best we can live together
Lukas Bogacz @lukasbogacz.com writing at loreley.one
Arjun Guha @arjunguha.bsky.social hacker / CS professor https://www.khoury.northeastern.edu/~arjunguha/
Laura Kopf @lkopf.bsky.social PhD student in Interpretable Machine Learning at @tuberlin.bsky.social & @bifold.berlin
https://web.ml.tu-berlin.de/author/laura-kopf/
claudia shi @claudiashi.bsky.social machine learning, causal inference, science of llm, ai safety, phd student @bleilab, keen bean
https://www.claudiashi.com/
Javier Ferrando @javifer.bsky.social Interpretability
@neelnanda.bsky.social @neelnanda.bsky.social
@woog0.bsky.social @woog0.bsky.social
Tim Hua @timhua.bsky.social Helping people is good I guess
Trying to do AI interp and control
Used to do economics
timhua.me
Koyena Pal @koyena.bsky.social CS Ph.D. Candidate @ Northeastern | Interpretability + Data Science | BS/MS @ Brown
koyenapal.github.io
Ben Stewart 🔸 @benstew.bsky.social AI Program Officer at Longview Philanthropy. Own views.
🔸 giving 10% of my lifetime income to effective charities via Giving What We Can
@yudkowsky.bsky.social @yudkowsky.bsky.social
Alexander Berger @albrgr.bsky.social CEO of Coefficient Giving
Julia Wise @juliawise.bsky.social Trying for human-compatible humans
Lynette Bye @lynettebye.bsky.social
Buck Shlegeris @bshlgrs.bsky.social
Vidur Kapur @vidurkapur.bsky.social Superforecaster at Good Judgment. Also forecasting at Swift Centre, Samotsvety, RAND and a hedge fund.
@jacobtref.bsky.social @jacobtref.bsky.social blog.jacobtrefethen.com
OpenAI Foundation
science!
Ryan Briggs @ryancbriggs.net Raising kids & bread & grant money. Cleaning data & diapers & fish. EA (bed nets, not light cone). Social scientist. typos. twitter.com/ryancbriggs
tetra 💎⏹️🇺🇳 @thetetra.space 💎 here to believe true things and do good actions 💎 someone should probably solve AI alignment 💎 enjoying things rules! ☀️ but it's not snowing now
english/toki pona/日本語
Carl Robichaud @carlrobi.bsky.social Program Officer on nuclear policy at Longview Philanthropy (http://longview.org). Opinions are my own.
Astral Codex Ten | Scott Alexander | Substack [Unofficial] @astralcodexten.com.web.brid.gy P(A|B) = [P(A)*P(B|A)]/P(B), all the rest is commentary. Click to read Astral Codex Ten, by Scott Alexander, a Substack publication.
🌉 bridged from 🌐 https://astralcodexten.com/: https://fed.brid.gy/web/astralcodexten.com
@wdmacaskill.bsky.social @wdmacaskill.bsky.social
Sebastian Farquhar @sebfar.bsky.social Senior Research Scientist at Google DeepMind. AGI Alignment researcher. Views my dog's.
Garrison Lovely @garrisonlovely.bsky.social Author of Obsolete: The AI Industry's Trillion Dollar Race to Replace You—and How to Stop It (Nation Books). https://orbooks.com/catalog/obsolete/
Bylines: NYT, Nature, BBC, Bloomberg, MIT Tech Review, TIME + others
Rossa O'Keeffe-O'Donovan @rossaokod.bsky.social Research @ Open Philanthropy. Formerly economist at GPI / Nuffield College, Oxford.
Interests: development econ, animal welfare, global catastrophic risks
Aaron Gertler @aarongertler.bsky.social Comms officer @ Open Philanthropy, former Magic pro, webfiction connoisseur. https://aarongertler.net/
Aaron Bergman @aaronbergman18.bsky.social 👎: suffering | 👍: EA, AI alignment, decoupling, R, cringe, amateur pharmacology + programming | Georgetown '22 (math+econ+phil) | Career status: 🤷♂️
@ozziegooen.bsky.social @ozziegooen.bsky.social
William Eden @weden.bsky.social
@dwarkesh.bsky.social @dwarkesh.bsky.social
Aaron Scher @aaronscher.bsky.social Technical AI Governance Research at MIRI
Views are my own
Nathan @handle.invalid Geopolitics, prediction markets. FLF Fellow. Capital case tweets are literal, others less. I like most people when I meet them.
Anka Reuel ➡️ NeurIPS @ankareuel.bsky.social Computer Science PhD Student @ Stanford | Geopolitics & Technology Fellow @ Harvard Kennedy School/Belfer | Vice Chair EU AI Code of Practice | Views are my own
Adam Binksmith @binksmith.com Building theaidigest.org and forecasting tools @aidigest.bsky.social
https://binksmith.com
Epoch AI @epochai.bsky.social We are a research institute investigating the trajectory of AI for the benefit of society.
epoch.ai
Yo Shavit @yonashav.bsky.social policy for v smart things @openai. Past: PhD @HarvardSEAS/@SchmidtFutures/@MIT_CSAIL. Posts my own; on my head be it
METR @metr.org METR is a research nonprofit that builds evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.
catherine 🌀 @catherinebrewer.bsky.social ai governance @openphil, unsupervised learner
Eli Lifland @elifland.bsky.social
Toby Ord @tobyord.bsky.social Senior Researcher at Oxford University.
Author — The Precipice: Existential Risk and the Future of Humanity.
tobyord.com
Trevor Levin @trevorlevin.bsky.social Trying to help the world navigate potentially transformative technologies, currently via AI Governance and Policy at Coefficient Giving. Enjoyer of acoustic guitars, history books, and plant-based foods.
Dean W. Ball @deanwb.bsky.social
https://hyperdimensional.co
harry law @harrylaw.bsky.social Thinking about thinking machines | University of Cambridge and Leverhulme Centre for the Future of Intelligence | Previously Google DeepMind
Richard Ngo @richardngo.bsky.social What would we need to understand in order to design an amazing future? Ex DeepMind, OpenAI
Samuel Hammond @hamandcheese.bsky.social Social policy synthesizer. www.secondbest.ca
allie lawsen @lexlawsen.bsky.social AI grantmaking at Coefficient Giving
Previously 80,000 Hours
lawsen.substack.com
Seb Krier @sebkrier.com friendly deep sea dweller
Surya Ganguli @suryaganguli.bsky.social Professor of Applied Physics at Stanford | Venture Partner a16z | Research in AI, Neuroscience, Physics
Guide Labs @guidelabs.bsky.social AI systems and models that are engineered to be interpretable and auditable.
www.guidelabs.ai
Ben Edelman @benedelman.bsky.social Thinking about how/why AI works/doesn't, and how to make it go well for us.
Currently: AI Agent Security @ US AI Safety Institute
benjaminedelman.com
Caden @cadentj.bsky.social
@joshengels.bsky.social @joshengels.bsky.social PhD student at MIT.
Working on mechanistic interpretability and AI safety.
wint @dril.bsky.social Never Bullshit
I challenge any and every one who wants to kick my ass to a debate .
https://www.patreon.com/dril
https://www.instagram.com/dril
https://linktr.ee/drilreal
Somin W @sominw.bsky.social cs phd @ northeastern.
Gonçalo Paulo @goncalo-paulo.bsky.social Interpretability researcher at @eleutherai.bsky.social
Ekdeep Singh @ ICML @ekdeepl.bsky.social Postdoc at CBS, Harvard University
(New around here)
Serena Booth @reniebird.bsky.social CS Prof at Brown University, PI of the GIRAFFE lab, former AI Policy Advisor in the US Senate, co-chair of the ACM Tech Policy Subcommittee on AI and Algorithms.
PhD at MIT CSAIL '23, Harvard '16, former Google APM. Dog mom to Ducki.
James Michaelov @jamichaelov.bsky.social Postdoc at Oxford. Research: language, the brain, NLP.
jmichaelov.com
Eran Malach @emalach.bsky.social Research Fellow @ Kempner Institute, Harvard University
Theory of Deep Learning / Learning of Deep Theory
Anna Tsvetkov @annatsv.bsky.social Postdoc @ Princeton AI
Natural and Artificial Minds
Prev: Philosophy PhD @ Brown, MIT FutureTech
Website: https://annatsv.github.io/
Zhaofeng Wu @zhaofengwu.bsky.social PhD student @ MIT | Previously PYI @ AI2 | MS'21 BS'19 BA'19 @ UW | zhaofengwu.github.io
Martin Wattenberg @wattenberg.bsky.social Human/AI interaction. ML interpretability. Visualization as design, science, art. Professor at Harvard, and part-time at Google DeepMind.
Sheridan Feucht @sfeucht.bsky.social PhD student doing LLM interpretability with @davidbau.bsky.social and @byron.bsky.social. (they/them) https://sfeucht.github.io
Imke Grabe @imkegrabe.bsky.social ☆ °。⋆ (mechanistic) interpretability + interaction (design) ⋆。° ☆
@eleutherai.bsky.social @eleutherai.bsky.social
Clément Dumas @butanium.bsky.social Master student at ENS Paris-Saclay / aspiring AI safety researcher / improviser
Prev research intern @ EPFL w/ wendlerc.bsky.social and Robert West
MATS Winter 7.0 Scholar w/ neelnanda.bsky.social
https://butanium.github.io
Aaron Mueller @amuuueller.bsky.social Postdoc at Northeastern and incoming Asst. Prof. at Boston U. Working on NLP, interpretability, causality. Previously: JHU, Meta, AWS
Mor Geva @megamor2.bsky.social https://mega002.github.io
Niklas Stoehr @niklasstoehr.bsky.social Gemini Post-Training ⚫️ Research Scientist at Google DeepMind ⚫️ PhD from ETH Zurich
Nina Rimsky @ninarimsky.bsky.social AI Safety Research // Software Engineering
Naomi Saphra @nsaphra.bsky.social Waiting on a robot body. All opinions are universal and held by both employers and family. ML/NLP professor.
nsaphra.net
Dashiell @dashiells.bsky.social Machine learning haruspex
Joe Stacey @joestacey.bsky.social NLP PhD student at Imperial College London and Apple AI/ML Scholar.
Sweta Karlekar @swetakar.bsky.social Machine learning PhD student @ Blei Lab in Columbia University
Working in mechanistic interpretability, nlp, causal inference, and probabilistic modeling!
Previously at Meta for ~3 years on the Bayesian Modeling & Generative AI teams.
🔗 www.sweta.dev
Nicolas Beltran-Velez @velezbeltran.bsky.social Machine Learning PhD Student
@ Blei Lab & Columbia University.
Working on probabilistic ML | uncertainty quantification | LLM interpretability.
Excited about everything ML, AI and engineering!
Daniel Johnson @ddjohnson.bsky.social PhD student at Vector Institute / University of Toronto. Building tools to study neural nets and find out what they know. He/him.
www.danieldjohnson.com
Alex Makelov @amakelov.bsky.social Mechanistic interpretability
Creator of https://github.com/amakelov/mandala
prev. Harvard/MIT
machine learning, theoretical computer science, competition math.
Andrew Lee @ajyl.bsky.social Post-doc @ Harvard. PhD UMich. Spent time at FAIR and MSR. ML/NLP/Interpretability
Martina G. Vilas @martinagvilas.bsky.social AI Evaluation & Interpretability @NVIDIA
Isabelle Lee @wordscompute.bsky.social ml/nlp phding @ usc, currently visiting harvard;
training & interpretability & reasoning
iglee.me
Pepa Atanasova @apepa.bsky.social Assistant Professor, University of Copenhagen; interpretability, xAI, factuality, accountability, xAI diagnostics https://apepa.github.io/
Federico Adolfi @fedeadolfi.bsky.social Computation & Complexity | AI Interpretability | Meta-theory | Computational Cognitive Science
Max Planck Institute ESI.
Dept. of Computer Science & Mathematics,
Goethe University Frankfurt.
https://fedeadolfi.github.io
On the job market!
Lee Sharkey @leesharkey.bsky.social Scruting matrices @ Apollo Research
Kayo Yin @kayoyin.bsky.social PhD student at UC Berkeley. NLP for signed languages and LLM interpretability. kayoyin.github.io
🏂🎹🚵♀️🥋
Julian Minder @jkminder.bsky.social PhD at EPFL with Robert West, Master at ETHZ
Mainly interested in Language Model Interpretability and Model Diffing.
MATS 7.0 Winter 2025 Scholar w/ Neel Nanda
jkminder.ch
Nishant Subramani @ ACL @nsubramani23.bsky.social PhD student @CMU LTI - working on model #interpretability, student researcher @google; prev predoc @ai2; intern @MSFT
nishantsubramani.github.io
Eric Todd @ericwtodd.bsky.social CS PhD Student, Northeastern University - Machine Learning, Interpretability https://ericwtodd.github.io