AI & ML interests
Data-centric AI, geospatial and tabular data, dataset curation, data provenance, and reproducible research.
Recent Activity
Axiom Codex
Axiom Codex turns public-source data into documented, schema-backed research datasets for employment, public health, places, and building permits. Maintained by Axiomancer Labs, the featured releases provide Parquet data, field definitions, source attribution, and machine-readable quality evidence. Read each dataset card for its exact scope, upstream terms, and limitations.
Full research releases
These larger source-record releases are distinct from the 25,000-row sample repositories below. Their exact upstream terms and any required attribution are repository-specific; this organization page does not assign a shared license or grant additional rights.
| Dataset | Contents and reported coverage | Repository |
|---|---|---|
| LEHD Workplace Area Characteristics (WAC), 2023 | 2,251,290 workplace-block records; 48 U.S. states plus D.C. (Alaska and Michigan are not included). This is WAC—not origin–destination commuter-flow data. LODES counts include Census disclosure-protection noise; small-area counts are approximate. | Axiom LEHD Workplace Jobs — national |
| CDC PLACES, 2025 tract data | 83,522 tract-level records across all 50 states and D.C. Modeled area estimates use 2023 tract geography and a mix of 2022 and 2023 survey years—not individual health observations. | Axiom CDC PLACES Tracts 2025 |
| POI Intelligence — California | 1,941,553 places from Overture release 2026-09-23.1. Static snapshot; names and point locations can describe sole traders or home-based businesses. Component license texts and attribution notices are included. |
Axiom POI Intelligence — California |
Public 25,000-row samples
These smaller, static samples make it possible to inspect the schemas and source-specific caveats before working with the larger releases. A sample is a separate artifact: its geography, dates, and source mix may differ from a larger release in the same family.
| Dataset | What the sample contains | Key limits |
|---|---|---|
| POI Intelligence sample | 25,000 places sampled from Overture Maps release 2026-09-23.1, with category, source identifiers, point location, and H3 resolution-8 cell. Global static snapshot. |
Coverage varies by region and category; it is thin for any one city. Missing operating status is not evidence that a place is open, and snapshot dates do not establish openings or closures. |
| LEHD Workplace Jobs (WAC) sample | 25,000 Census LODES8 WAC workplace-block records for 2023, sampled across California, Florida, Illinois, New York, Texas, and Washington. | WAC is not origin–destination flow data. It covers six states and one year; noise-infused small-block counts are approximate, not exact job counts. |
| Permit Signals sample | 25,000 permit records from Chicago, Los Angeles, New York City, and Seattle. | City periods differ (the Los Angeles feed ends in May 2023); status is a collection-time snapshot, not a verified outcome. Fields are missing at different rates by city, and rows without reviewed mappings are excluded. |
Quick start
Install the Hugging Face Datasets library, then load a sample by repository ID:
pip install datasets
from datasets import load_dataset
places = load_dataset(
"AxiomCodex/axiom-poi-intelligence-sample-25k",
split="train",
)
print(places[0])
The repository cards include dataset-specific schema tables, citations, source terms, and links to the release manifest and quality evidence where provided. See the POI sample card, WAC sample card, and permit sample card before using their fields.
Intended use and limits
- These are source-record datasets for research, exploration, and analysis of the observations described in each repository card. They are not supervised-learning benchmarks: this page reports no benchmark scores, ground-truth evaluation labels, or train/validation/test evaluation design.
- The featured repositories use a single Hugging Face
trainsplit as a transport/tooling convention, not an ML evaluation split. Their cards retain the Axiom Codex governed training-use boundary; a split name is not training permission. - No organization-wide license applies. Read each repository's own license metadata and card, follow the upstream source terms and attribution notices that apply to its records, and do not transfer one dataset's rights to another. This README grants no additional training or commercial rights.
- These research repositories and samples are not commercial packages. Any separately offered package has its own scope and terms; a public sample is not a substitute for that package or its license.
- Source coverage, dates, missingness, bias, and record meaning differ by dataset. A normalized field or status value does not by itself make a source observation complete, current, independently verified, or suitable for a decision about a person or property.
Browse the interactive catalog and public sample viewer, public standards and crosswalks, or Axiom Codex website.
Questions and corrections
For questions, card corrections, or removal requests, open a discussion in the Axiom Codex README Space community or the relevant dataset's Community tab. Include the repository ID, revision, and affected field or record identifier; do not post personal information or credentials.
For business inquiries, use enterprise@axiomcodex.io, the contact published on the Axiom Codex website. Dataset questions and corrections belong in the Community discussions above.