227 GB
25 files
Updated about 1 month ago
Name
Size
.gitattributes2.5 kB
xet
README.md3.59 kB
xet
build_index.py1.43 kB
xet
idx_aadhar.0.parquet16.9 GB
xet
idx_aadhar.1.parquet15.7 GB
xet
idx_aadhar.2.parquet15.7 GB
xet
idx_aadhar.3.parquet15.7 GB
xet
idx_aadhar.4.parquet21.3 GB
xet
idx_aadhar.5.parquet26.1 GB
xet
idx_aadhar.6.parquet11.7 GB
xet
idx_phone.0.parquet17.1 GB
xet
idx_phone.1.parquet16.7 GB
xet
idx_phone.2.parquet16.6 GB
xet
idx_phone.3.parquet17.1 GB
xet
idx_phone.4.parquet16.9 GB
xet
idx_phone.5.parquet16.5 GB
xet
idx_phone.6.parquet3.21 GB
xet
main.py13.5 kB
xet
requirements.txt55 Bytes
xet
run_api.bat251 Bytes
xet
setup.bat2.71 kB
xet
setup.ps13.04 kB
xet
setup.sh2.55 kB
xet
split_idx.py3.37 kB
xet
upload_parts.py2.04 kB
xet
README.md

ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup

Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).

Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).

Contents

The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP upload path caps files at 50 GB. The API reads all parts of an index with one glob — zone-map pruning still makes lookups milliseconds-fast.

Files Total What
idx_phone.0.parquet … idx_phone.6.parquet ~102 GB Data sorted by phoneNumber → phone lookups ~1s
idx_aadhar.0.parquet … idx_aadhar.6.parquet ~100 GB Data sorted by aadharNumber → aadhar lookups ~2s
main.py — FastAPI app (DuckDB-backed, 15-way parallel, dedup max 2)
build_index.py — Rebuild indexes from raw parquet (only if you ever need to)
split_idx.py — Split a built index into < 50 GB parts
setup.bat — One-click restore (Windows)
setup.ps1 — One-click restore (PowerShell)
setup.sh — One-click restore (Linux / VPS)

Raw data files (part1.parquet, part2a.parquet, part2b_new.parquet, 104 GB) live in the Icmr-and-hitek repo — the setup scripts download them from there automatically.

One-click restore

Pick your platform, run ONE file. It downloads data + indexes + code, creates a venv, installs deps, starts the API.

Windows: double-click setup.bat PowerShell: right-click → Run with PowerShell, or powershell -ExecutionPolicy Bypass -File setup.ps1 Linux/VPS:

chmod +x setup.sh
./setup.sh

Total download: ~305 GB (104 GB data + ~200 GB index parts). Scripts resume interrupted downloads (curl -C -), so a dropped connection is not a problem — just re-run.

API

Base: http://127.0.0.1:8001 (Linux: 0.0.0.0:8001)

Endpoint Use
GET /search?q=<phone> Phone search — ~1s (indexed)
GET /search?q=<aadhar> Aadhar search — ~2s (indexed)
GET /search?q=<name>&limit=10 Name/text search (falls back to raw scan, slower)
GET /search?q=X&field=district&mode=exact Single-field search
POST /search/parallel Batch: up to 50 searches in parallel
GET /health Status incl. which indexes are active
GET /docs Interactive Swagger UI

All 11 columns searchable: name, fathersName, phoneNumber, aadharNumber, otherNumber, address, district, pincode, state, town, source. Duplicates capped at 2 per person. source = icmr | inddata (hitek data is labelled inddata).

Manual start (if scripts already ran)

# Windows
cd /d "C:\path\to\folder"
start "" .venv\Scripts\python.exe -m uvicorn main:app --host 127.0.0.1 --port 8001

# Linux
cd /path/to/folder
ICMR_DATA_DIR="$PWD/data" nohup .venv/bin/python -m uvicorn main:app --host 0.0.0.0 --port 8001 > api.log 2>&1 &

ICMR_DATA_DIR points at the data/ folder; if it is not set, data/ next to main.py is used. main.py auto-detects index parts (idx_phone.*.parquet) — no config needed.

Rebuilding indexes (rarely needed)

ICMR_DATA_DIR="$PWD/data" .venv/bin/python build_index.py

Each index takes 6-11 hours on a typical machine. After building, split_idx.py can split them into upload-sized parts again.

Total size
227 GB
Files
25
Last updated
Aug 28
Pre-warmed CDN
US EU US EU

Contributors