DevToM: Developmental Theory of Mind Benchmark for LLMs

Python (Inspect) · R · Summer 2026

Problem
LLM benchmarks usually report one aggregate score, which can't show which reasoning skills a model has or whether they emerge in a human-like order.
Data
201 items across 12 Theory of Mind dimensions, each anchored to empirically established child age norms (ages 2.5 to 11); 5,000+ graded responses from 28 models across 5 providers, drawn from a 53-model routing panel.
Method
Evaluation harness in Inspect (AISI) with custom datasets, solvers, and scorers; items red-teamed against contamination and answer leakage; dual analytic pipelines combining mastery-based age mapping with mixed-effects and IRT models that include task-format covariates.
Result
Model profiles are non-monotonic across all 12 dimensions, with sizeable age-equivalent swings between adjacent skills. Results are explorable in the Streamlit dashboard below, served from a Parquet snapshot.
LLM Benchmarking IRT & Mixed-Effects Models Streamlit Dashboard AI Evaluation

Family Goals

TypeScript (Next.js) · SwiftUI · 2026

Problem
Families had no shared place to set and track daily and weekly health goals together.
Data
Water, exercise, and step logs per family member, with steps synced automatically from iOS HealthKit.
Method
Full-stack web and iOS app on a SQLite-compatible database (Turso), with password-protected family accounts and assignable goals.
Result
Deployed web and iOS apps with weekly, monthly, and yearly progress views, from completion rates to calendar heatmaps.
Full-Stack SQL (SQLite) Mobile

Infant Egocentric Object Recognition

Python (PyTorch, Detectron2) · 2020–2021

Problem
Which objects do infants actually see day to day, and can standard object detectors handle the unusual angles of head-mounted camera footage?
Data
10,000 annotated frames of infant egocentric video recorded in home environments.
Method
Fine-tuned and tested Detectron2 object detection and segmentation models, evaluating performance on infant-view image angles.
Result
While models initially are really poor at identifying objects from the infant viewpoint, minimal training improved segmentation and recognition in a few categories.
Computer Vision Model Fine-Tuning Model Evaluation

Peekbank: Multi-Lab Eye-Tracking Data Standardization

R (Tidyverse) · 2020–2021

Problem
Infant looking-while-listening studies come from many labs, each with its own formats and protocols, which made cross-study comparison impractical.
Data
Eye-tracking datasets contributed by 12+ labs, including 5 datasets I imported.
Method
Built the import and validation layer: defined the relational schema, wrote ETL scripts for 5 contributed datasets, and added automated schema checks that flag protocol mismatches at ingestion.
Result
A shared open-access database enabling cross-study analyses of early word recognition.
ETL Data Validation Relational Schema Open Data

Gaze Entropy & Decision-Making (LDM Analysis)

Jupyter Notebook · 2022

Problem
Does how people spread their gaze during a reinforcement-learning task relate to the choices they make?
Data
Eye-tracking data recorded during a reinforcement-learning decision task.
Method
Pipeline for processing raw eye-tracking data, computing gaze entropy, and correlating it with decision-making outcomes.
Eye-Tracking Data Pipelines Reinforcement Learning

Pregnancy During COVID-19: Brainhack Project

Jupyter / Neuroimaging · 2026

Problem
How does pregnancy during the COVID-19 pandemic relate to children's later neurobehavioural development?
Data
Functional and structural MRI connectivity data.
Method
Functional and structural connectivity analysis pipeline, built collaboratively during a Brainhack event.
MRI Analysis Reproducible Pipelines