Michal Stefanik @michalstefanik.bsky.social NLP Researcher: Robust Language Models | Masaryk University & GaussAlgo
Clément Dumas @butanium.bsky.social Master student at ENS Paris-Saclay / aspiring AI safety researcher / improviser
Prev research intern @ EPFL w/ wendlerc.bsky.social and Robert West
MATS Winter 7.0 Scholar w/ neelnanda.bsky.social
https://butanium.github.io
Mor Geva @megamor2.bsky.social https://mega002.github.io
Niklas Stoehr @niklasstoehr.bsky.social Gemini Post-Training ⚫️ Research Scientist at Google DeepMind ⚫️ PhD from ETH Zurich
Nina Rimsky @ninarimsky.bsky.social AI Safety Research // Software Engineering
Dashiell @dashiells.bsky.social Machine learning haruspex
Nicolas Beltran-Velez @velezbeltran.bsky.social Machine Learning PhD Student
@ Blei Lab & Columbia University.
Working on probabilistic ML | uncertainty quantification | LLM interpretability.
Excited about everything ML, AI and engineering!
Andrew Lee @ajyl.bsky.social Post-doc @ Harvard. PhD UMich. Spent time at FAIR and MSR. ML/NLP/Interpretability
Lee Sharkey @leesharkey.bsky.social Scruting matrices @ Apollo Research
Arthur Conmy @arthurconmy.bsky.social Aspiring 10x reverse engineer at Google DeepMind
Eric Todd @ericwtodd.bsky.social CS PhD Student, Northeastern University - Machine Learning, Interpretability https://ericwtodd.github.io
vedang @vedanglad.bsky.social ai interpretability research and running • thinking about how models think • prev @MIT cs + physics
Jonathan Ling @jonling.bsky.social Assistant Professor @HopkinsMedicine @JHUPath
https://scholar.google.com/citations?user=dGBD72YAAAAJ
Cristina @cristinaml.bsky.social ML/AI researcher @JohnsHopkins
Shan Chen @shan23chen.bsky.social PhDing @Harvard @MassGenBrigham|PhD Fellow @Google | Previously @Bos_CHIP @BrandeisU
More robustness and explainabilities 🧐 for Health AI.
shanchen.dev
Jannik Brinkmann @jannikbrinkmann.bsky.social
Francesco Ortu @francescortu.bsky.social NLP & Interpretability | PhD Student @ University of Trieste & Laboratory of Data Engineering of Area Science Park | Prev MPI-IS
Carl Allen @carl-allen.bsky.social Laplace Junior Chair, Machine Learning
ENS Paris. (prev ETH Zurich, Edinburgh, Oxford..)
Working on mathematical foundations/probabilistic interpretability of ML (what NNs learn🤷♂️, disentanglement🤔, king-man+woman=queen?👌…)
@michael-pearce.bsky.social @michael-pearce.bsky.social
Dilyara Bareeva @dilya.bsky.social PhD Candidate in Interpretability @FraunhoferHHI | 📍Berlin, Germany
dilyabareeva.github.io
Alessandro Stolfo @alestolfo.bsky.social PhD @ ETHZ - LLM Interpretability
alestolfo.github.io
Bart Bussmann @bartbussmann.bsky.social Independent Mechanistic Interpretability Researcher
@kevdududu.bsky.social @kevdududu.bsky.social
Koyena Pal @koyena.bsky.social CS Ph.D. Candidate @ Northeastern | Interpretability + Data Science | BS/MS @ Brown
koyenapal.github.io
@woog0.bsky.social @woog0.bsky.social
@neelnanda.bsky.social @neelnanda.bsky.social
Javier Ferrando @javifer.bsky.social Interpretability
Taufeeque @taufeeque.bsky.social Research Engineer @ FAR.AI
taufeeque9.github.io
Gonçalo Paulo @goncalo-paulo.bsky.social Interpretability researcher at @eleutherai.bsky.social
Kunvar Thaman @firstuserhere.bsky.social
dribnet @drib.net creations with code and networks
Caden @cadentj.bsky.social
@joshengels.bsky.social @joshengels.bsky.social PhD student at MIT.
Working on mechanistic interpretability and AI safety.
Ana Lučić @a-lucic.bsky.social Assistant professor at the University of Amsterdam. Previously at Microsoft Research, Partnership on AI.
Yoann Poupart @xmaster6y.bsky.social XAI PhD Student & Entrepreneur
claudia shi @claudiashi.bsky.social machine learning, causal inference, science of llm, ai safety, phd student @bleilab, keen bean
https://www.claudiashi.com/
NDIF Team @ndif-team.bsky.social The National Deep Inference Fabric, an NSF-funded computational infrastructure to enable research on large-scale Artificial Intelligence.
🔗 NDIF: https://ndif.us
🧰 NNsight API: https://nnsight.net
😸 GitHub: https://github.com/ndif-team/nnsight
Hidenori Tanaka @hidenori8tanaka.bsky.social Group Leader, CBS-NTT "Physics of Intelligence" Program at Harvard
website: https://sites.google.com/view/htanaka/home
Can @canrager.bsky.social
Actionable Interpretability Workshop ICML2025 @actinterp.bsky.social 🛠️ Actionable Interpretability🔎 @icmlconf.bsky.social 2025 | Bridging the gap between insights and actions ✨ https://actionable-interpretability.github.io
Hadas Orgad @hadasorgad.bsky.social
Sheridan Feucht @sfeucht.bsky.social PhD student doing LLM interpretability with @davidbau.bsky.social and @byron.bsky.social. (they/them) https://sfeucht.github.io
Antonin Poché @antoninpoche.bsky.social PhD Student doing XAI for NLP at @ANITI_Toulouse, IRIT, and IRT Saint Exupery.
🛠️ Interpreto & Xplique library development team member.
https://antoninpoche.github.io/
Arnab Sen Sharma @arnabsensharma.bsky.social PhD Student at Northeastern, working to make LLMs interpretable
Oliver Eberle @eberleoliver.bsky.social Senior Researcher Machine Learning at BIFOLD | TU Berlin 🇩🇪
Prev at IPAM | UCLA | BCCN
Interpretability | XAI | NLP & Humanities | ML for Science
Sarah Wiegreffe @sarah-nlp.bsky.social Research in NLP (mostly LM interpretability & explainability).
Assistant prof at UMD CS + CLIP.
Previously @ai2.bsky.social @uwnlp.bsky.social
Views my own.
sarahwie.github.io
Thomas Fel @thomasfel.bsky.social Explainability, Computer Vision, Neuro-AI.🪴 Kempner Fellow @Harvard.
Prev. PhD @Brown, @Google, @GoPro. Crêpe lover.
📍 Boston | 🔗 thomasfel.me
José Oramas @jaom7.bsky.social Associate Professor @UAntwerp, sqIRL/IDLab, imec.
#RepresentationLearning, #Model #Interpretability & #Explainability
A guy who plays with toy bricks, enjoys research and gaming.
Opinions are my own
idlab.uantwerpen.be/~joramasmogrovejo
Kirill Bykov @kirillbykov.bsky.social PhD student in Interpretable ML @UMI_Lab_AI, @bifoldberlin, @TUBerlin
Nils Feldhus @nfel.bsky.social Postdoctoral Researcher @ University of Groningen / GroNLP (@gronlp.bsky.social) interested in the interpretability and analysis of language models. Ex- BIFOLD, TU Berlin, DFKI. https://nfelnlp.github.io/
Zeynep Akata @zeynepakata.bsky.social Liesel Beckmann Distinguished Professor of Computer Science at Technical University of Munich and Director of the Institute for Explainable ML at Helmholtz Munich
Zining Zhu @zhuzining.bsky.social Asst Prof @ Stevens. Working on NLP, Explainable, Safe and Trustworthy AI. https://ziningzhu.github.io
Chris Olah @colah.bsky.social Reverse engineering neural networks at Anthropic. Previously Distill, OpenAI, Google Brain.Personal account.
Andreas Madsen @andreasmadsen.bsky.social Ph.D. in NLP Interpretability from Mila. Previously: independent researcher, freelancer in ML, and Node.js core developer.
David Atkinson @diatkinson.bsky.social PhD student at Northeastern, previously at EpochAI. Doing AI interpretability.
diatkinson.github.io
Qianli Wang @ ACL 2025🇦🇹 @qiaw99.bsky.social Second-year PhD student at XplaiNLP group @TU Berlin: interpretability & explainability
Website: https://qiaw99.github.io
Daniel Johnson @ddjohnson.bsky.social PhD student at Vector Institute / University of Toronto. Building tools to study neural nets and find out what they know. He/him.
www.danieldjohnson.com
Alex Makelov @amakelov.bsky.social Mechanistic interpretability
Creator of https://github.com/amakelov/mandala
prev. Harvard/MIT
machine learning, theoretical computer science, competition math.
Xiaoyan Bai @elenal3ai.bsky.social PhD @UChicagoCS / BE in CS @Umich / ✨AI/NLP transparency and interpretability/📷🎨photography painting
Gunnar König @gunnark.bsky.social PostDoc @ Uni Tübingen
explainable AI, causality
gunnarkoenig.com
Christoph Molnar @christophmolnar.bsky.social Author of Interpretable Machine Learning and other books
Newsletter: https://mindfulmodeler.substack.com/
Website: https://christophmolnar.com/
Berk Ustun @berkustun.bsky.social Assistant Prof at UCSD. I work on safety, interpretability, and fairness in machine learning. www.berkustun.com
Chris Wendler @wendlerc.bsky.social Postdoc at the interpretable deep learning lab at Northeastern University, deep learning, LLMs, mechanistic interpretability
Katharina Beckh @kbeckh.bsky.social Data Scientist at Fraunhofer IAIS
PhD Student at University of Bonn
Lamarr Institute
XAI, NLP, Human-centered AI
Lorenz Linhardt @lorenzlinhardt.bsky.social PhD Student at the TU Berlin ML group + BIFOLD | BUA Fellow
Model robustness/correction 🤖🔧
Understanding representation spaces 🌌✨
Fiona K. Ewald @fionaewald.bsky.social PhD Student @ LMU Munich
Munich Center for Machine Learning (MCML)
Research in Interpretable ML / Explainable AI
Thiago Serra @thserra.bsky.social Assistant professor at University of Iowa, formerly at Bucknell University, mathematical optimizer with an #orms PhD from Carnegie Mellon University, curious about scaling up constraint learning, proud father of two
Ribana Roscher @ribana.bsky.social Professor of Machine Learning in Agriculture at University of Bonn
Working on Explainable ML🔍, Data-centric ML🐿️, Sustainable Agriculture🌾, Earth Observation Data Analysis🌍, and more...
André Panisson @panisson.bsky.social Principal Researcher @ CENTAI.eu | Leading the Responsible AI Team. Building Responsible AI through Explainable AI, Fairness, and Transparency. Researching Graph Machine Learning, Data Science, and Complex Systems to understand collective human behavior.
Harry Cheon @scheon.com "Seung Hyun" | MS CS & BS Applied Math @UCSD 🌊 | LPCUWC 18' 🇭🇰 | AI Evaluation, Safety, Alignment | 🇰🇷
harry.scheon.com
Julian Minder @jkminder.bsky.social PhD at EPFL with Robert West, Master at ETHZ
Mainly interested in Language Model Interpretability and Model Diffing.
MATS 7.0 Winter 2025 Scholar w/ Neel Nanda
jkminder.ch
Simon Schrodi @simonschrodi.bsky.social 🎓 PhD student @cvisionfreiburg.bsky.social @UniFreiburg
💡 interested in mechanistic interpretability, robustness, AutoML & ML for climate science
https://simonschrodi.github.io/
Sebastian Bordt @sbordt.bsky.social Language models and interpretable machine learning. Postdoc @ Uni Tübingen.
https://sbordt.github.io/
Martin Wattenberg @wattenberg.bsky.social Human/AI interaction. ML interpretability. Visualization as design, science, art. Professor at Harvard, and part-time at Google DeepMind.
Daniel Arp @darpsky.bsky.social Assistant Professor at TU Wien
Machine Learning & Security
Farnoush Rezaei-Jafari @farnoushrj.bsky.social ML Ph.D. Candidate @tuberlin.bsky.social and @bifold.berlin | Explainable AI, Interpretability, Efficient Machine Learning
farnoushrj.github.io
Stefan Haufe @sparsity.bsky.social Professor of Machine Learning at TUBerlin, group leader at PTB. Lab account: @qailabs.bsky.social.
@sparsity@mastodon.social
tu.berlin/uniml/about/head-of-group
Johanna Vielhaben @johannavielhaben.bsky.social PhD candidate in the XAI group at Fraunhofer HHI
Sophia Sanborn @naturecomputes.bsky.social Searching for principles of neural representation | Neuro + AI @ enigmaproject.ai | Stanford | sophiasanborn.com
Ruizhe Li @ruizheli.bsky.social Assistant Professor at University of Aberdeen | Postdoc at UCL | PhD at University of Sheffield | mechanistic interpretability & multimodal LLMs | https://www.ruizhe.space
Conor O'Sullivan @conorosullyds.bsky.social PhD in progress - XAI and ML for coastal monitoring 🌊
My content: linktr.ee/conorosullyds
@dheerajr.bsky.social @dheerajr.bsky.social
Laura Kopf @lkopf.bsky.social PhD student in Interpretable Machine Learning at @tuberlin.bsky.social & @bifold.berlin
https://web.ml.tu-berlin.de/author/laura-kopf/
Max Lamparth, Ph.D. @mlamparth.bsky.social Research Fellow @ Stanford Intelligent Systems Laboratory and Hoover Institution at Stanford University | Focusing on interpretable, safe, and ethical AI/LLM decision-making. Ph.D. from TUM.
Anka Reuel ➡️ NeurIPS @ankareuel.bsky.social Computer Science PhD Student @ Stanford | Geopolitics & Technology Fellow @ Harvard Kennedy School/Belfer | Vice Chair EU AI Code of Practice | Views are my own
Sukrut Rao @sukrutrao.bsky.social PhD Student at the Max Planck Institute for Informatics @cvml.mpi-inf.mpg.de @maxplanck.de | Explainable AI, Computer Vision, Neuroexplicit Models
Web: sukrutrao.github.io
Guilherme Alves @asilvaguilherme.bsky.social Machine Learning • Fairness • Explainable ML • AutoML
http://guilhermealves.eti.br
Kanishka Misra @kanishka.bsky.social Assistant Professor of Linguistics at UT Austin. Works on computational understanding of language, concepts, and generalization. Aspiring wugologist!
🕸️👁️: https://kanishka.website
Kyle Mahowald @kmahowald.bsky.social UT Austin linguist http://mahowak.github.io/. computational linguistics, cognition, psycholinguistics, NLP, crosswords. occasionally hockey?
Sebastian Schuster @sebschu.bsky.social Computational semantics and pragmatics, interpretability and occasionally some psycholinguistics. he/him. 🦝
https://sebschu.com
Aaron Mueller @amuuueller.bsky.social Postdoc at Northeastern and incoming Asst. Prof. at Boston U. Working on NLP, interpretability, causality. Previously: JHU, Meta, AWS
Nathan Schneider @complingy.bsky.social Computational Linguist and Professional Nerd at Georgetown University
he/him pronouns, ALL the prepositions. http://nathan.cl
Roger Levy @rplevy.bsky.social Director, MIT Computational Psycholinguistics Lab. President, Cognitive Science Society. Chair of the MIT Faculty. Open access & open science advocate. He.
Lab webpage: http://cpl.mit.edu/
Personal webpage: https://www.mit.edu/~rplevy
Tal Linzen @tallinzen.bsky.social NYU professor, Google research scientist. Good at LaTeX.
Tiago Pimentel @tpimentel.bsky.social Postdoc at ETH. Formerly, PhD student at the University of Cambridge :)
David Mortensen @davidrmortensen.bsky.social I make colorless green GPUs sleep brrriously. Computational phonology, morphology, language change models, speech/language technologies (especially for people with disabilities).
Tom McCoy @rtommccoy.bsky.social Assistant professor at Yale Linguistics. Studying computational linguistics, cognitive science, and AI. He/him.
Valentin Hofmann @valentinhofmann.bsky.social Assistant Professor @cislmu.bsky.social @lmu.de
Chris Potts @cgpotts.bsky.social Stanford Professor of Linguistics and, by courtesy, of Computer Science, and member of @stanfordnlp.bsky.social and The Stanford AI Lab. He/Him/His. https://web.stanford.edu/~cgpotts/
Aryaman Arora @aryaman.io member of technical staff @stanfordnlp.bsky.social