Selected Projects
Projects
Data pipelines, statistical models, and evaluation systems, each laid out as problem → data → method → result. See more on GitHub.
DevToM: Developmental Theory of Mind Benchmark for LLMs
- Problem
- LLM benchmarks usually report one aggregate score, which can't show which reasoning skills a model has or whether they emerge in a human-like order.
- Data
- 201 items across 12 Theory of Mind dimensions, each anchored to empirically established child age norms (ages 2.5 to 11); 5,000+ graded responses from 28 models across 5 providers, drawn from a 53-model routing panel.
- Method
- Evaluation harness in Inspect (AISI) with custom datasets, solvers, and scorers; items red-teamed against contamination and answer leakage; dual analytic pipelines combining mastery-based age mapping with mixed-effects and IRT models that include task-format covariates.
- Result
- Model profiles are non-monotonic across all 12 dimensions, with sizeable age-equivalent swings between adjacent skills. Results are explorable in the Streamlit dashboard below, served from a Parquet snapshot.
LLM Benchmarking
IRT & Mixed-Effects Models
Streamlit Dashboard
AI Evaluation
Family Goals
- Problem
- Families had no shared place to set and track daily and weekly health goals together.
- Data
- Water, exercise, and step logs per family member, with steps synced automatically from iOS HealthKit.
- Method
- Full-stack web and iOS app on a SQLite-compatible database (Turso), with password-protected family accounts and assignable goals.
- Result
- Deployed web and iOS apps with weekly, monthly, and yearly progress views, from completion rates to calendar heatmaps.
Full-Stack
SQL (SQLite)
Mobile
Infant Egocentric Object Recognition
- Problem
- Which objects do infants actually see day to day, and can standard object detectors handle the unusual angles of head-mounted camera footage?
- Data
- 10,000 annotated frames of infant egocentric video recorded in home environments.
- Method
- Fine-tuned and tested Detectron2 object detection and segmentation models, evaluating performance on infant-view image angles.
- Result
- While models initially are really poor at identifying objects from the infant viewpoint, minimal training improved segmentation and recognition in a few categories.
Computer Vision
Model Fine-Tuning
Model Evaluation
Peekbank: Multi-Lab Eye-Tracking Data Standardization
- Problem
- Infant looking-while-listening studies come from many labs, each with its own formats and protocols, which made cross-study comparison impractical.
- Data
- Eye-tracking datasets contributed by 12+ labs, including 5 datasets I imported.
- Method
- Built the import and validation layer: defined the relational schema, wrote ETL scripts for 5 contributed datasets, and added automated schema checks that flag protocol mismatches at ingestion.
- Result
- A shared open-access database enabling cross-study analyses of early word recognition.
ETL
Data Validation
Relational Schema
Open Data
Gaze Entropy & Decision-Making (LDM Analysis)
- Problem
- Does how people spread their gaze during a reinforcement-learning task relate to the choices they make?
- Data
- Eye-tracking data recorded during a reinforcement-learning decision task.
- Method
- Pipeline for processing raw eye-tracking data, computing gaze entropy, and correlating it with decision-making outcomes.
Eye-Tracking
Data Pipelines
Reinforcement Learning
Pregnancy During COVID-19: Brainhack Project
- Problem
- How does pregnancy during the COVID-19 pandemic relate to children's later neurobehavioural development?
- Data
- Functional and structural MRI connectivity data.
- Method
- Functional and structural connectivity analysis pipeline, built collaboratively during a Brainhack event.
MRI Analysis
Reproducible Pipelines