Data Science · AI Evaluation · Cognitive Science

Naiti S. Bhatt

Data scientist and research engineer building validated data pipelines, statistical models, and AI evaluation systems, informed by developmental cognitive science and open science practices. I am currently focused on how insights from childhood learning and cognitive mechanisms can help us evaluate and improve AI systems.

Portrait of Naiti S. Bhatt
Focus Areas

What I work on

The common thread is measurement and data quality: getting messy human and model data into a trustworthy shape, then checking that what we measure is what we meant to measure.

Data Pipelines & Engineering

Designing scalable, reproducible ETL and analysis pipelines for multi-source human-subjects data (MRI, eye-tracking, behavioral questionnaires, egocentric video), with schema definitions and automated validation checks at ingestion so problems are caught before they reach analysis.

Statistical Modeling & Validation

Fitting mixed-effects, regression, and IRT models to noisy, longitudinal human data, with an emphasis on validity: whether a measure actually captures what it claims to.

LLM Evaluation & Alignment

Evaluating large language models using tasks validated for benchmarking human learning and cognition across domains of social reasoning.

Developmental Cognitive Science

Modeling neurocognitive processes underlying skill (object understanding, decision-making, and social reasoning) development across childhood.

Background

A bit more about me

I completed my masters at the Department of Psychology at the University of Edinburgh, where I built reproducible analysis pipelines for functional and diffusion MRI to study the neural mechanisms of theory of mind development for my dissertation work. Before Edinburgh, I graduated from Scripps College with majors in Computer Science and Neuroscience, where I trained and tested computer vision models to segment and detect objects in infant egocentric video frames to characterize early object learning for my thesis work. Most recently, I built DevToM, a developmental Theory of Mind benchmark for large language models, and I evaluate frontier models at Outlier & Alignerr.

I'm open to data science, research engineering, and AI evaluation roles in the NYC area. Get in touch or view my CV.

I grew up in Basking Ridge, New Jersey. Outside of research, I roast and brew specialty coffee, chase sunrises and sunsets, and swim in international waters (19 countries & 30 US states so far, from 4°C to 40°C).