Rashid Kisejjere
Data scientist at Pollicy
Machine Learning Engineer specializing in low-resource NLP and production ML systems for African languages. Over
3 years of hands-on experience deploying end-to-end ML solutions in healthcare and accessibility domains. Strong
track record in multilingual model development, RAG systems, and MLOps with 3 peer-reviewed publications.
Abstract
Artificial General Intelligence (AGI), increasingly feels less like science fiction and more like an approaching reality. Models are already writing code, generating speech, solving research problems, and assisting in scientific discovery. Yet despite these capabilities, we still do not have a clear understanding of how modern AI systems actually work internally. We can train them, scale them, and deploy them, but their learned algorithms remain largely opaque. Mechanistic interpretability aims to change that. Rather than treating neural networks as black boxes, this field seeks to reverse-engineer the circuits, representations, and computations inside models. What features do neurons detect? How do attention heads collaborate? What algorithms emerge during training? And how can we systematically uncover them? This session provides a practical introduction to mechanistic interpretability. We will explore the core concepts behind circuits and internal representations, and examine simple hands-on techniques for probing models. The goal is not just to understand the tools, but to understand why interpretability is becoming critical in a world where AGI may be around the corner. If we are building increasingly powerful systems, understanding their internal mechanics is no longer optional. This talk equips participants with the conceptual foundation and practical starting point needed to begin opening up modern neural networks and investigating what they are really doing under the hood.