AI & ML interests

Education, Tutoring

Recent Activity

sushijoe  updated a dataset about 2 months ago
NationalTutoringObservatory/MathEd-PII
sushijoe  updated a Space about 2 months ago
NationalTutoringObservatory/README
sushijoe  published a Space about 2 months ago
NationalTutoringObservatory/README
View all activity

Organization Card

National Tutoring Observatory

National Tutoring Observatory

Our mission is to improve teaching and learning at scale by learning from great tutors.

We are a team of researchers, educators, and technologists building the world's largest dataset of authentic tutoring interactions, along with the open-source research infrastructure needed to work with it responsibly. We partner equitably with tutoring providers, school districts, and underserved communities.

🌐 nationaltutoringobservatory.org · 💻 GitHub

What we work on

  • Tutoring data at scale — observing real tutor–student interactions across subjects and grade levels to understand what effective tutoring actually looks like.
  • Open-source research infrastructure — tools for annotation, de-identification, and multimodal processing of tutoring data (chat, video, transcripts, whiteboard activity).
  • Privacy-preserving data release — methods for sharing educational dialogue that protect the students and teachers in it while keeping the data useful for research.

Sandpiper

Sandpiper is our open-source AI textual-annotation application (MIT licensed). It powers the annotation and de-identification workflows behind our datasets, including the proposer–reviewer pipeline we use to detect and replace personally identifiable information in tutoring transcripts.

Research

Our work on utility-preserving de-identification examines a problem specific to math tutoring: numeric expressions often look like structured identifiers, so naive redaction destroys the instructional content it is meant to protect.

📄 Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset — Zhou, Vanacore, Ahtisham, Lee, Pietrzak, Hedley, Dias, Shaw, Schäfer, Kizilcec (arXiv:2602.16571)

The paper introduces MathEd-PII, a companion benchmark dataset for PII detection in math tutoring dialogue.

A note on access

Our datasets contain authentic educational interactions, so we release them under gated access: the dataset card and terms are readable by anyone, while downloading requires agreeing to our Data Use Terms and a short review of intended use. This lets us share data that would otherwise stay locked away, while keeping commitments to the students, tutors, and institutions who made it possible.

Datasets that are not currently listed here are in preparation or under review. Please get in touch if you have questions about availability.

Where our data has been de-identified using a hide-in-plain-sight approach, identifiers you see in the text are surrogates — realistic but fabricated substitutes. Dataset cards explain this in detail; please read them before use.

Contact

General inquiries and data access questions: sandpiper_admin@cornell.edu

Acknowledgments

This work is supported by the National Science Foundation (Grant No. 2321499), the Gates Foundation, and the Chan Zuckerberg Initiative. Any opinions, findings, and conclusions are those of the authors and do not necessarily reflect the views of the funders.

models 0

None public yet