Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| data | 10 items | ||
| .gitattributes | 2.5 kB xet | 738f1125 | |
| README.md | 3.62 kB xet | b92f266d |
SpecMine
A large-scale corpus of Spec-Driven Development (SDD) artifacts mined from public GitHub: the Markdown specifications that AI coding agents turn into code, attributed to the tool that produced them, enriched with repository metadata and full commit history, and, for a subsample of repositories, followed into the pull requests that change them.
This is the per-table Parquet mirror of SpecMine v1.0 (July 2026 snapshot). The citable archive of record (with the MySQL dump, CSV, and spec-text JSONL) is on Zenodo; the schema, loaders, an example notebook, and a 500-repository sample are on GitHub at https://github.com/shyamagarwal13/specmine-official. SpecMine is the dataset for the MSR 2027 Mining Challenge.
Tables (select a config above)
| Config | Rows | What it is |
|---|---|---|
spec_files |
470,795 | broad-census spine (repo object, provenance, tool attribution) |
spec_file_commits |
780,335 | per-file commit history |
spec_content_features |
468,307 | 39 structural / requirement-template features per spec |
spec_links |
2,421,323 | traceability index (typed spec→code references) |
kiro_files |
98,574 | Kiro requirements/design/tasks (metadata) |
openspec_artifact_files |
266,230 | OpenSpec proposals/designs/task lists (+ content) |
openspec_code_refs |
435,401 | OpenSpec task→code refs resolved on the git tree |
pull_requests |
5,992 | spec-touching PRs (11 tools unioned, 10 non-empty), tool column |
pr_files |
348,141 | per-file changesets, each file flagged spec/code |
repo_trees |
73,030 | git-tree metadata references resolve against |
Spec text is not on Hugging Face; it ships as specs.jsonl.gz in the Zenodo
record and joins to spec_files on file_url_sha16.
Usage
from datasets import load_dataset
prs = load_dataset("USER/specmine", "pull_requests", split="train")
Or read a table directly:
import pandas as pd
prs = pd.read_parquet("hf://datasets/USER/specmine/data/pull_requests.parquet")
License and ethics
Compilation released under CC BY 4.0. Underlying file content remains under each
source repository's own license (see repo_license_spdx_id). Commit emails,
avatar URLs, and GPG signatures are removed; public logins and numeric IDs are
kept. Please do not attempt deanonymization.
Citation
@inproceedings{specmine2027,
author = {Agarwal, Shyam and Vasilescu, Bogdan},
title = {SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts},
year = {2027},
booktitle = {Proceedings of the 24th International Conference on Mining Software Repositories},
series = {MSR '27}
}
- Total size
- 1.6 GB
- Files
- 12
- Last updated
- Aug 31
- Pre-warmed CDN
- US EU US EU