AI & ML interests
Education, Tutoring
Recent Activity
National Tutoring Observatory
Our mission is to improve teaching and learning at scale by learning from great tutors.
We are a team of researchers, educators, and technologists building the world's largest dataset of authentic tutoring interactions, along with the open-source research infrastructure needed to work with it responsibly. We partner equitably with tutoring providers, school districts, and underserved communities.
🌐 nationaltutoringobservatory.org · 💻 GitHub
What we work on
- Tutoring data at scale — observing real tutor–student interactions across subjects and grade levels to understand what effective tutoring actually looks like.
- Open-source research infrastructure — tools for annotation, de-identification, and multimodal processing of tutoring data (chat, video, transcripts, whiteboard activity).
- Privacy-preserving data release — methods for sharing educational dialogue that protect the students and teachers in it while keeping the data useful for research.
Sandpiper
Sandpiper is our open-source AI textual-annotation application (MIT licensed). It powers the annotation and de-identification workflows behind our datasets, including the proposer–reviewer pipeline we use to detect and replace personally identifiable information in tutoring transcripts.
Research
Our work on utility-preserving de-identification examines a problem specific to math tutoring: numeric expressions often look like structured identifiers, so naive redaction destroys the instructional content it is meant to protect.
📄 Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset — Zhou, Vanacore, Ahtisham, Lee, Pietrzak, Hedley, Dias, Shaw, Schäfer, Kizilcec (arXiv:2602.16571)
The paper introduces MathEd-PII, a companion benchmark dataset for PII detection in math tutoring dialogue.
A note on access
Our datasets contain authentic educational interactions, so we release them under gated access: the dataset card and terms are readable by anyone, while downloading requires agreeing to our Data Use Terms and a short review of intended use. This lets us share data that would otherwise stay locked away, while keeping commitments to the students, tutors, and institutions who made it possible.
Datasets that are not currently listed here are in preparation or under review. Please get in touch if you have questions about availability.
Where our data has been de-identified using a hide-in-plain-sight approach, identifiers you see in the text are surrogates — realistic but fabricated substitutes. Dataset cards explain this in detail; please read them before use.
Contact
General inquiries and data access questions: sandpiper_admin@cornell.edu
Acknowledgments
This work is supported by the National Science Foundation (Grant No. 2321499), the Gates Foundation, and the Chan Zuckerberg Initiative. Any opinions, findings, and conclusions are those of the authors and do not necessarily reflect the views of the funders.