Walter Laurito @walterlaurito.bsky.social
Geodesic Research @geodesicresearch.bsky.social We're behind http://alignmentpretraining.ai.
Let's align some AIs.
https://geodesicresearch.ai/
Wyatt Walls @wwalls.bsky.social Tech lawyer. Generates plausible bullshit in 6 minute increments. More active on https://x.com/lefthanddraft
Jérôme Guyot @jerome-gyt.bsky.social Ignoring obvious arguments for aesthetic reasons.
Quantum, cryptography & complexity · ENS Paris-Saclay
@ben300694.bsky.social @ben300694.bsky.social
Arnab Sen Sharma @arnabsensharma.bsky.social PhD Student at Northeastern, working to make LLMs interpretable
Explainable AI Berlin @xai-berlin.bsky.social Explainable AI research from the machine learning group of Prof. Klaus-Robert Müller at @tuberlin.bsky.social & @bifold.berlin
Deniz Bayazit @bayazitdeniz.bsky.social #NLProc PhD student @EPFL
#interpretability
Amir Zur @amirzur.bsky.social PhD @stanfordnlp.bsky.social
Verona Teo @veronateo.bsky.social
Psyho @psyho.bsky.social Game Designer; Problem Solver; past: OpenAI (Dota), Pro Competitive Programmer, Poker
Maxime Méloux @maximemeloux.bsky.social PhD student @LIG | Causal abstraction, interpretability & LLMs
David Duvenaud @davidduvenaud.bsky.social Machine learning prof at U Toronto. Working on evals and AGI governance.
Antonin Poché @antoninpoche.bsky.social PhD Student doing XAI for NLP at @ANITI_Toulouse, IRIT, and IRT Saint Exupery.
🛠️ Interpreto & Xplique library development team member.
https://antoninpoche.github.io/
Kerem Sahin @keremsahin22.bsky.social MS CS @ Northeastern
Byron Wallace @byron.bsky.social Assoc. Prof in CS @ Northeastern, NLP/ML & health & etc. He/him.
Hadas Orgad @hadasorgad.bsky.social
Actionable Interpretability Workshop ICML2025 @actinterp.bsky.social 🛠️ Actionable Interpretability🔎 @icmlconf.bsky.social 2025 | Bridging the gap between insights and actions ✨ https://actionable-interpretability.github.io
Steve Byrnes @stevebyrnes.bsky.social Researching Artificial General Intelligence Safety, via thinking about neuroscience and algorithms, at Astera Institute. https://sjbyrnes.com/agi.html
@aranguri.bsky.social @aranguri.bsky.social
Martin Wattenberg @wattenberg.bsky.social Human/AI interaction. ML interpretability. Visualization as design, science, art. Professor at Harvard, and part-time at Google DeepMind.
@dimkakha.bsky.social @dimkakha.bsky.social
Can @canrager.bsky.social
Cas (Stephen Casper) @scasper.bsky.social Computer scientist working on AI safeguards, incidents, & governance research. Assistant professor @harvardkennedy.bsky.social @harvard.edu.
https://stephencasper.com/
almost dribnet @dribnet.bsky.social I moved here -> https://bsky.app/profile/drib.net <- here moved I
@vkrakovna.bsky.social @vkrakovna.bsky.social Research scientist in AI alignment at Google DeepMind. Co-founder of Future of Life Institute. Views are my own and do not represent GDM or FLI.
Shreyans @pyparrot.bsky.social Interpretability, AI ethics, Reinforcement Learning
Sara Fish @sarafish.bsky.social PhD student at Harvard interested in EconCS and ML / previously Caltech undergrad in math
Hidenori Tanaka @hidenori8tanaka.bsky.social Group Leader, CBS-NTT "Physics of Intelligence" Program at Harvard
website: https://sites.google.com/view/htanaka/home
Core Francisco Parkg @corefpark.bsky.social https://cfpark00.github.io/
Ido Aizenbud @idoai.bsky.social Computational Neuroscience PhD Student
@evhub.bsky.social @evhub.bsky.social Alignment Stress-Testing Team Lead at Anthropic. Opinions my own. Previously: MIRI, OpenAI, Google, Yelp, Ripple. (he/him/his)
Aengus Lynch @aengusl.bsky.social AI safety researcher
Tim Hua @timhua.bsky.social Helping people is good I guess
Trying to do AI interp and control
Used to do economics
timhua.me
@yimingliu.bsky.social @yimingliu.bsky.social
Praneet @praneet.bsky.social ML PhD at McGill
Dennis Fucci @dennisfucci.bsky.social Speech | XAI | Fairness in AI
PhD student @fbk-mt.bsky.social
@emmabortz.bsky.social @emmabortz.bsky.social
Angie Boggust @angieboggust.bsky.social MIT PhD candidate in the VIS group working on interpretability and human-AI alignment
Simon Schrodi @simonschrodi.bsky.social 🎓 PhD student @cvisionfreiburg.bsky.social @UniFreiburg
💡 interested in mechanistic interpretability, robustness, AutoML & ML for climate science
https://simonschrodi.github.io/
Sarah Wiegreffe @sarah-nlp.bsky.social Research in NLP (mostly LM interpretability & explainability).
Assistant prof at UMD CS + CLIP.
Previously @ai2.bsky.social @uwnlp.bsky.social
Views my own.
sarahwie.github.io
Patrick Kahardipraja @pkhdipraja.bsky.social PhD student @ Fraunhofer HHI. Interpretability, incremental NLP, and NLU. https://pkhdipraja.github.io/
Sophie Hao @cinnamonlab.ai Assistant professor of Linguistics and Data Science at Boston University. NLP, computational linguistics, interpretability, social bias and fairness. she/her. https://www.notaphonologist.com/
Jason Lee @jasondeanlee.bsky.social Associate Professor at Princeton
Machine Learning Researcher
Sophia Sanborn @naturecomputes.bsky.social Searching for principles of neural representation | Neuro + AI @ enigmaproject.ai | Stanford | sophiasanborn.com
Jakub Łucki @jakublucki.bsky.social Visiting Researcher at NASA JPL | Data Science MSc at ETH Zurich
Max Lamparth, Ph.D. @mlamparth.bsky.social Research Fellow @ Stanford Intelligent Systems Laboratory and Hoover Institution at Stanford University | Focusing on interpretable, safe, and ethical AI/LLM decision-making. Ph.D. from TUM.
NDIF Team @ndif-team.bsky.social The National Deep Inference Fabric, an NSF-funded computational infrastructure to enable research on large-scale Artificial Intelligence.
🔗 NDIF: https://ndif.us
🧰 NNsight API: https://nnsight.net
😸 GitHub: https://github.com/ndif-team/nnsight
Laura Kopf @lkopf.bsky.social PhD student in Interpretable Machine Learning at @tuberlin.bsky.social & @bifold.berlin
https://web.ml.tu-berlin.de/author/laura-kopf/
Shu Yang @shuyhere.bsky.social CS phd @KAUST Building things
@danielchtan.bsky.social @danielchtan.bsky.social
Yuli Slavutsky @yulislavutsky.bsky.social Stats Postdoc at Columbia, @bleilab.bsky.social
Statistical ML, Generalization, Uncertainty, Empirical Bayes
https://yulisl.github.io/
claudia shi @claudiashi.bsky.social machine learning, causal inference, science of llm, ai safety, phd student @bleilab, keen bean
https://www.claudiashi.com/
@beodu.bsky.social @beodu.bsky.social
Yoann Poupart @xmaster6y.bsky.social XAI PhD Student & Entrepreneur
sqIRL Lab @sqirllab.bsky.social We are sqIRL(squirrel), the Interpretable Representation Learning Lab based at IDLab - University of Antwerp & imec.
Research Areas: #RepresentationLearning, #Interpretability, #explainability
#ML #AI #XAI #mechinterp
Website: https://sqirllab.github.io/
Andrea Santilli @asantilli.bsky.social PhD student in NLP at Sapienza | Prev: Apple MLR, @colt-upf.bsky.social , HF Bigscience, PiSchool, HumanCentricArt #NLProc
www.santilli.xyz
Jonas Rohweder @handle.invalid
Fabian Grob @fabiangrob.bsky.social CS @ TUM | relAI MSc Fellow
Ana Lučić @a-lucic.bsky.social Assistant professor at the University of Amsterdam. Previously at Microsoft Research, Partnership on AI. - A attentionmech @attentionmech.bsky.social
Yifei Wang @ywren.bsky.social
Paul @notpaulmartin.bsky.social NLP PhD @ Cambridge Language Technology Lab
paulsbitsandbytes.com
@joshengels.bsky.social @joshengels.bsky.social PhD student at MIT.
Working on mechanistic interpretability and AI safety.
Caden @cadentj.bsky.social
Ruizhe Li @ruizheli.bsky.social Assistant Professor at University of Aberdeen | Postdoc at UCL | PhD at University of Sheffield | mechanistic interpretability & multimodal LLMs | https://www.ruizhe.space
Solal Nathan @solalnathan.com PhD student @ U. Paris-Saclay / Inria, AI for social good, fairness, RecSys, congestion avoidance, optimal transport. ENS PS 2018.
Free software advocate, linux user, cat owner. asso @ auro.re and crans.org. Bicyle, Bouldering, improv.
solalnathan.com
Laura @lauraruis.bsky.social PhD supervised by Tim Rocktäschel and Ed Grefenstette, part time at Cohere. Language and LLMs. Spent time at FAIR, Google, and NYU (with Brenden Lake). She/her.
Andreas Madsen @andreasmadsen.bsky.social Ph.D. in NLP Interpretability from Mila. Previously: independent researcher, freelancer in ML, and Node.js core developer.
Tim Davidson @imtd.bsky.social 🌐 https://www.trdavidson.com
🔬research: deep generative learning; agentic systems; synthetic data
PhD @EPFL on reliable magic
Spent time @MSR, @Google
machine learning & company building
🎓@NYU @UvA alumn
dribnet @drib.net creations with code and networks
@jamesaoldfield.bsky.social @jamesaoldfield.bsky.social Visiting scholar @ UW-Madison & PhD student in machine learning @ QMUL. Interested in interpretability and AI safety.
https://james-oldfield.github.io/
Simon Lermen @simonlermen.bsky.social I work on AI safety and AI in cybersecurity
Sohee Yang @soheeyang.bsky.social PhD student/research scientist intern at UCL NLP/Google DeepMind (50/50 split). Previously MS at KAIST AI and research engineer at Naver Clova. #NLP #ML 👉 https://soheeyang.github.io/
Kunvar Thaman @firstuserhere.bsky.social
Gonçalo Paulo @goncalo-paulo.bsky.social Interpretability researcher at @eleutherai.bsky.social
Sekh (Sk) Mainul Islam @sekh-copenlu.bsky.social PhD Fellow at the CopeNLU Group, University of Copenhagen; working on explainable automatic fact-checking . Prev: NYU Abu Dhabi, IIT Kharagpur.
https://mainuliitkgp.github.io/
Yu-Min Tseng @ymtseng.bsky.social Master Student @NTU_TW | Visiting Student @UVA | Seeking 25 fall CS PhD 🎯
🏠 www.ymtseng.com
Taufeeque @taufeeque.bsky.social Research Engineer @ FAR.AI
taufeeque9.github.io
Tal Haklay @talhaklay.bsky.social NLP | Interpretability | PhD student at the Technion
Javier Ferrando @javifer.bsky.social Interpretability
Yorguin-Jose Mantilla-Ramos @yjmantilla.bsky.social veneco trying to get into interpretability, both for natural and artificial intelligence.
currently a masters student at Université de Montréal.
Rob Bensinger @robbensinger.bsky.social - C Andrew Critch @critch.bsky.social Human being. Trying to do good. CEO @ Encultured AI. AI Researcher @ UC Berkeley. Listed bday is approximate ;)
Alex Irpan @alexirpan.bsky.social Research Scientist @ Google DeepMind. Formerly Robotics, now AI Safety. Has a blog. Views are my own.
Eric Neyman @ericneyman.bsky.social Professional reference class tennis player. I like non-fillet frozen fish, packaged medicaments, and other oily seeds.
Samuel Albanie @samuelalbanie.bsky.social
Elliott Thornley @elliottthornley.bsky.social Research Fellow at Oxford University's Global Priorities Institute.
Working on the philosophy of AI.
@neelnanda.bsky.social @neelnanda.bsky.social
akbir khan @akbir.bsky.social dumbest overseer at @anthropic
https://www.akbir.dev
Alex Turner @turntrout.bsky.social Research scientist at Google DeepMind. All opinions are my own. Vegan, 10% of my income pledged to effective charities (GWWC)
https://turntrout.com
Tom Everitt @tom4everitt.bsky.social AGI safety researcher at Google DeepMind, leading causalincentives.com
Personal website: tomeveritt.se
Buck Shlegeris @bshlgrs.bsky.social
Epoch AI @epochai.bsky.social We are a research institute investigating the trajectory of AI for the benefit of society.
epoch.ai
David Lindner @davidlindner.bsky.social Making AI safer at Google DeepMind
davidlindner.me
Manoel Horta Ribeiro @manoelhortaribeiro.bsky.social Assistant Professor @ Princeton
Previously: EPFL 🇨🇭, UFMG 🇧🇷
Interests: Computational Social Science, Platforms, GenAI, Moderation
@cranchesco.bsky.social @cranchesco.bsky.social
Jean Kaddour @jeankaddour.bsky.social https://github.com/PySpur-Dev/PySpur
PhD Student at UCL // LLMs
Rifo Genadi @rifoag.bsky.social M.Sc. Student at MBZUAI. I just started doing Mech Interp. I also do some stuff on low-resource language research.