1.6 GB
12 files
Updated about 1 month ago
Name
Size
data
.gitattributes2.5 kB
xet
README.md3.62 kB
xet
README.md

SpecMine

A large-scale corpus of Spec-Driven Development (SDD) artifacts mined from public GitHub: the Markdown specifications that AI coding agents turn into code, attributed to the tool that produced them, enriched with repository metadata and full commit history, and, for a subsample of repositories, followed into the pull requests that change them.

This is the per-table Parquet mirror of SpecMine v1.0 (July 2026 snapshot). The citable archive of record (with the MySQL dump, CSV, and spec-text JSONL) is on Zenodo; the schema, loaders, an example notebook, and a 500-repository sample are on GitHub at https://github.com/shyamagarwal13/specmine-official. SpecMine is the dataset for the MSR 2027 Mining Challenge.

Tables (select a config above)

Config Rows What it is
spec_files 470,795 broad-census spine (repo object, provenance, tool attribution)
spec_file_commits 780,335 per-file commit history
spec_content_features 468,307 39 structural / requirement-template features per spec
spec_links 2,421,323 traceability index (typed spec→code references)
kiro_files 98,574 Kiro requirements/design/tasks (metadata)
openspec_artifact_files 266,230 OpenSpec proposals/designs/task lists (+ content)
openspec_code_refs 435,401 OpenSpec task→code refs resolved on the git tree
pull_requests 5,992 spec-touching PRs (11 tools unioned, 10 non-empty), tool column
pr_files 348,141 per-file changesets, each file flagged spec/code
repo_trees 73,030 git-tree metadata references resolve against

Spec text is not on Hugging Face; it ships as specs.jsonl.gz in the Zenodo record and joins to spec_files on file_url_sha16.

Usage

from datasets import load_dataset
prs = load_dataset("USER/specmine", "pull_requests", split="train")

Or read a table directly:

import pandas as pd
prs = pd.read_parquet("hf://datasets/USER/specmine/data/pull_requests.parquet")

License and ethics

Compilation released under CC BY 4.0. Underlying file content remains under each source repository's own license (see repo_license_spdx_id). Commit emails, avatar URLs, and GPG signatures are removed; public logins and numeric IDs are kept. Please do not attempt deanonymization.

Citation

@inproceedings{specmine2027,
  author    = {Agarwal, Shyam and Vasilescu, Bogdan},
  title     = {SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts},
  year      = {2027},
  booktitle = {Proceedings of the 24th International Conference on Mining Software Repositories},
  series    = {MSR '27}
}
Total size
1.6 GB
Files
12
Last updated
Aug 31
Pre-warmed CDN
US EU US EU

Contributors