hashundefined/full / README.md
hashundefined's picture
|
download
raw
3.59 kB
---
license: mit
---
# ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the [Kzr0xx/Icmr-and-hitek](https://huggingface.co/datasets/Kzr0xx/Icmr-and-hitek) dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took **~17 hours** of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
## Contents
The indexes are stored as **sorted parts** (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP upload path caps files at 50 GB. The API reads all parts of an index with one glob — zone-map pruning still makes lookups milliseconds-fast.
| Files | Total | What |
|---|---|---|
| `idx_phone.0.parquet` … `idx_phone.6.parquet` | ~102 GB | Data sorted by phoneNumber → phone lookups ~1s |
| `idx_aadhar.0.parquet` … `idx_aadhar.6.parquet` | ~100 GB | Data sorted by aadharNumber → aadhar lookups ~2s |
| `main.py` | — | FastAPI app (DuckDB-backed, 15-way parallel, dedup max 2) |
| `build_index.py` | — | Rebuild indexes from raw parquet (only if you ever need to) |
| `split_idx.py` | — | Split a built index into < 50 GB parts |
| `setup.bat` | — | One-click restore (Windows) |
| `setup.ps1` | — | One-click restore (PowerShell) |
| `setup.sh` | — | One-click restore (Linux / VPS) |
Raw data files (`part1.parquet`, `part2a.parquet`, `part2b_new.parquet`, 104 GB) live in the **Icmr-and-hitek** repo — the setup scripts download them from there automatically.
## One-click restore
Pick your platform, run ONE file. It downloads data + indexes + code, creates a venv, installs deps, starts the API.
**Windows:** double-click `setup.bat`
**PowerShell:** right-click → Run with PowerShell, or `powershell -ExecutionPolicy Bypass -File setup.ps1`
**Linux/VPS:**
```bash
chmod +x setup.sh
./setup.sh
```
Total download: ~305 GB (104 GB data + ~200 GB index parts). Scripts resume interrupted downloads (`curl -C -`), so a dropped connection is not a problem — just re-run.
## API
Base: `http://127.0.0.1:8001` (Linux: `0.0.0.0:8001`)
| Endpoint | Use |
|---|---|
| `GET /search?q=<phone>` | Phone search — **~1s** (indexed) |
| `GET /search?q=<aadhar>` | Aadhar search — **~2s** (indexed) |
| `GET /search?q=<name>&limit=10` | Name/text search (falls back to raw scan, slower) |
| `GET /search?q=X&field=district&mode=exact` | Single-field search |
| `POST /search/parallel` | Batch: up to 50 searches in parallel |
| `GET /health` | Status incl. which indexes are active |
| `GET /docs` | Interactive Swagger UI |
All 11 columns searchable: name, fathersName, phoneNumber, aadharNumber, otherNumber, address, district, pincode, state, town, source. Duplicates capped at 2 per person. `source` = `icmr` | `inddata` (hitek data is labelled inddata).
## Manual start (if scripts already ran)
```bash
# Windows
cd /d "C:\path\to\folder"
start "" .venv\Scripts\python.exe -m uvicorn main:app --host 127.0.0.1 --port 8001
# Linux
cd /path/to/folder
ICMR_DATA_DIR="$PWD/data" nohup .venv/bin/python -m uvicorn main:app --host 0.0.0.0 --port 8001 > api.log 2>&1 &
```
`ICMR_DATA_DIR` points at the `data/` folder; if it is not set, `data/` next to `main.py` is used. `main.py` auto-detects index parts (`idx_phone.*.parquet`) — no config needed.
## Rebuilding indexes (rarely needed)
```bash
ICMR_DATA_DIR="$PWD/data" .venv/bin/python build_index.py
```
Each index takes 6-11 hours on a typical machine. After building, `split_idx.py` can split them into upload-sized parts again.

Xet Storage Details

Size:
3.59 kB
·
Xet hash:
375881198e29dce679eb2f83cae9c13e43cd63b561291fc679427a1cae4e6bf2

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.