Saatvik Kher
PhD Candidate in Computer Science, University of California, Irvine
Advised by Padhraic Smyth
I develop statistical methods that make large language model evaluations reliable, well-calibrated, and cost-efficient.
About
I am a 3rd-year Computer Science PhD candidate at the University of California, Irvine, where I am fortunate to be advised by Padhraic Smyth.
I develop statistical methods for uncertainty quantification and calibration in large language models, with a focus on making LLM evaluations and LLM-as-a-judge scores reliable, well-calibrated, and cost-efficient. My work combines Bayesian inference with sequential decision-making to quantify uncertainty in black-box model behavior and to adaptively allocate evaluation and labeling effort. I also study online learning with multi-armed bandits for human-AI interaction and algorithmic fairness.
At UC Irvine I am a data curator for the UCI Machine Learning Repository and a member of the Steckler Center for Responsible, Ethical, and Accessible Technology. Since May 2026 I have also been a research collaborator at Google.
Before Irvine I earned a B.A. in Computer Science and Mathematics from Pomona College. I held ML research positions with the SMALL NSF REU at Williams College and at the Yale School of Medicine.
Education
-
2024 – 2028 (expected)
Ph.D. Candidate in Computer Science
University of California, Irvine · Advisor: Padhraic Smyth -
2024 – 2026
M.S. in Computer Science
University of California, Irvine -
2020 – 2024
B.A. in Computer Science & Mathematics, cum laude
Pomona College
News
- Started a research collaboration with Google, working on LLM calibration and uncertainty quantification.
- “Measuring the Impact of Missingness in Traffic Stop Data” published in Harvard Data Science Review.
- “Bayesian Evaluation of Large Language Model Behavior” appeared at the NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle.
- Teaching assistant for CS177: Applications of Probability in Computer Science at UC Irvine.
- Began my PhD at UC Irvine, supported by a Computer Science Department Research Fellowship.
Publications
* denotes joint first authorship. See also my Google Scholar profile.
Published
-
Measuring the Impact of Missingness in Traffic Stop Data
Harvard Data Science Review, January 2026 -
Bayesian Evaluation of Large Language Model Behavior
NeurIPS 2025 Workshop on Evaluating the Evolving LLM LifecycleWorkshop arXiv -
How to Use Causal Inference to Study Use of Force
CHANCE, 37(4), 2024Magazine DOI -
Identifying Game-Based Digital Biomarkers of Cognitive Risk for Adolescent Substance Misuse: Protocol for a Proof-of-Concept Study
JMIR Research Protocols, November 2023Journal DOI -
A Machine Learning Model Using In-Game Data for Predicting Unhealthy Substance Use Among Adolescents
MLHC Abstracts, 2023Conference PDF
Under review
-
Calibrated Prediction of Ordinal Scores with Distributional LLM Judges
Preprint
-
Sequential Bayesian Evaluation of Large Language Model Behavior
Preprint
-
Bayesian Learning to Select Fair Classifiers
Preprint
-
Learning to Assign Prediction Tasks to Agents with Capacity Constraints
Preprint arXiv
- Improving and Evaluating Machine Learning Methods for Forensic Shoeprint Matching
Research Experience
-
May 2026 – present
Research Collaborator
Google, Irvine, CA- Building calibration and uncertainty quantification methods for LLMs, anomaly detection for large-scale time series, and skill routing for LLM agents.
-
Jan 2025 – present
Graduate Student Researcher
UC Irvine- Developed Bayesian and distributional methods for uncertainty quantification in LLM evaluation: calibrated LLM-as-a-judge models that capture annotator disagreement on ordinal rating tasks, and a sequential Bayesian framework that quantifies stochasticity in benchmark scores and adaptively selects prompts for cost-effective evaluation.
- Designed Bayesian multi-armed bandit algorithms for online selection and assignment of human and AI agents, jointly optimizing predictive accuracy, subgroup fairness, and capacity constraints, with regret guarantees and significant empirical gains.
-
May 2022 – Dec 2024
Research Fellow
Pomona College- Conducted a statistical analysis of missing-data mechanisms and racial bias in 200M+ US traffic stops from the Stanford Open Policing Project.
- Applied causal inference methods to estimate racial disparities in police use of force, addressing selection bias and non-random missingness in observational data.
-
Jun 2023 – Aug 2023
REU Researcher
SMALL (NSF REU), Williams College- Improved robustness of machine learning methods for forensic shoeprint matching under distribution shift using novel clustering and phase-correlation similarity features.
- Built and deployed an explainable machine learning application for forensic shoeprint matching, surfacing model evidence to forensic practitioners.
-
Jun 2022 – Sep 2022
Machine Learning Researcher
Yale University School of Medicine- Analyzed video game log data to identify features predictive of substance misuse in teens.
- Trained and evaluated interpretable machine learning models in Sklearn for predicting unhealthy substance use, supporting early intervention.
Fellowships & Awards
- Sep 2024UCI Computer Science Department Research Fellowship
- May 2024Sigma Xi
- Oct 2023Best Poster Award, NESS-NextGen Data Science Day
- May 2023Kenneth Cooke Summer Research Fellowship
- May 2022Pomona College Summer Undergraduate Research Project (SURP)
- 2020 – 2024Pomona College Scholar
Teaching
- CS175: Projects in AI (NLP)UC IrvineWinter 2026
- CS177: Applications of Probability in Computer ScienceUC IrvineFall 2025
- CS140: AlgorithmsPomona CollegeSpring 2024
- CS181SY: Managing Complex SystemsPomona CollegeSpring 2024
- CS158: Machine LearningPomona CollegeFall 2023
- MATH158: Statistical Linear ModelsPomona CollegeSpring 2023
- CS054: Discrete Math & Functional ProgrammingPomona CollegeFall 2022
Skills
- Programming languages
- Python, Java, R, SQL
- ML libraries & tools
- PyTorch, Sklearn, pandas, Docker, Git
- Methods
- Bayesian inference, uncertainty quantification, calibration, LLM evaluation, online learning, time series analysis