YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- Prefered System :: NVIDIA RTX A4000 | NVIDIA RTX 3090 Ti
- π³ Docker (recommended β run this project on any machine)
- β‘ Quick Start + π Miniconda & Conda Environment + π Setup NGROK
- π¦ Python Packages
- Install all dependencies
- π Audio Backend
- π§ Patch face_recognition_models
- π Secrets & Credentials (not tracked in git)
- π HuggingFace Hub Login & Token
- π₯ Fix CUDA Library Path for GPU (cublas/cudnn)
Prefered System :: NVIDIA RTX A4000 | NVIDIA RTX 3090 Ti
Avoid Servers
- RTX 5060 Ti
- Tesla P100
π³ Docker (recommended β run this project on any machine)
docker-compose.yml runs standalone on any machine with plain Docker
Desktop (CPU-only β torch/ctranslate2/YOLO all fall back automatically,
just slower). GPU acceleration is opt-in via docker-compose.gpu.yml,
which requires the NVIDIA Container Toolkit on the host and a driver
supporting CUDA 13 forward compat (>=580.x). Everything below is a
one-time host setup β the image itself is fully self-contained and pinned.
# 1. Pull the actual model weights (image does NOT bake these in β 5GB, LFS)
git lfs pull
# 2. Provide secrets (never baked into the image β see .env.example)
cp .env.example tools/.env # then fill in real values
# also place: tools/service_account.json, modules/ekyc/utils/cloudVisionAPI.json
# 3. Build + run β CPU-only, any machine:
docker compose up --build -d
# ...or with GPU, on a host with the NVIDIA Container Toolkit installed:
docker compose -f docker-compose.yml -f docker-compose.gpu.yml up --build -d
# 4. Check it's alive
curl http://localhost:5656/docs
docker compose logs -f
To stop: docker compose down (add -v to also wipe the model-download
caches in speechbrain_cache/hf_cache, forcing a re-download next start).
Verifying the Docker setup
bash scripts/lint-docker.sh # hadolint + compose config, seconds
bash scripts/docker-smoke-test.sh # full build + start + wait-for-healthy
docker-smoke-test.sh needs the same prerequisites as step 1-2 above (real
AI_Models/ weights + credential files) since the app won't start without
them β that's also why CI (.github/workflows/docker-ci.yml) only lints and
builds, not runs, the image.
Running with plain docker run (no compose, e.g. on a machine that only pulled the image)
Compose's volumes/env/GPU flags translated 1:1:
docker run -d \
--name he-universe \
--gpus all \
-p 5656:5656 \
-e PUID=$(id -u) -e PGID=$(id -g) \
-v "$(pwd)/AI_Models:/app/AI_Models:ro" \
-v speechbrain_cache:/app/AI_Models/SpeechAnalyzer_Models/pretrained_models \
-v hf_cache:/home/appuser/.cache/huggingface \
-v "$(pwd)/logs:/app/logs" \
-v "$(pwd)/tools/.env:/app/tools/.env:ro" \
-v "$(pwd)/tools/service_account.json:/app/tools/service_account.json:ro" \
-v "$(pwd)/modules/ekyc/utils/cloudVisionAPI.json:/app/modules/ekyc/utils/cloudVisionAPI.json:ro" \
--restart unless-stopped \
hawkeyes-universe:latest
(swap the final image name for yourdockerhubuser/hawkeyes-universe:latest if pulling from Docker Hub instead of building locally)
Publishing to Docker Hub
The image is already built to be push-safe: AI_Models/ (proprietary weights),
all credential files, .git, and the Dockerfile/compose files themselves are
excluded from the build context (.dockerignore) β only application code and
pinned dependencies get baked in. Application source code IS baked in
though (COPY . .), so unless this is meant to be public, push to a
private Docker Hub repository.
export IMAGE_NAME=yourdockerhubuser/hawkeyes-universe
export IMAGE_TAG=1.0.0 # or "latest" β pick something you can track releases by
docker login
docker compose build
docker compose push
(docker compose push works because image: in docker-compose.yml is
already registry-qualified via $IMAGE_NAME/$IMAGE_TAG.)
To pull and run it on another machine afterward, that machine needs docker login
too (for a private repo), then the same docker run/docker compose up
commands above with IMAGE_NAME/IMAGE_TAG set to match, plus its own copy of
AI_Models/ (git-lfs) and the secret files β those never travel with the image.
β‘ Quick Start + π Miniconda & Conda Environment + π Setup NGROK
(manual/bare-metal setup β skip this if you're using Docker above)
ο ο¦ dev
sudo apt update && sudo apt upgrade -y
sudo apt install -y iproute2 libgl1 nano wget unzip nvtop git git-lfs build-essential cmake \
libopenblas-dev liblapack-dev libx11-dev libgtk-3-dev libglib2.0-0
git config --global credential.helper store
git clone -b dev https://huggingface.co/HawkEyesAI/v2_HE_Universe
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
chmod +x Miniconda3-latest-Linux-x86_64.sh
./Miniconda3-latest-Linux-x86_64.sh -b -p $HOME/miniconda3
export PATH="$HOME/miniconda3/bin:$PATH"
conda init
source ~/.bashrc
conda create --name HE python=3.11 -y
conda activate HE
curl -s https://ngrok-agent.s3.amazonaws.com/ngrok.asc | sudo tee /etc/apt/trusted.gpg.d/ngrok.asc >/dev/null
echo "deb https://ngrok-agent.s3.amazonaws.com buster main" | sudo tee /etc/apt/sources.list.d/ngrok.list
sudo apt update
sudo apt install -y ngrok
ngrok config add-authtoken "${NGROK_AUTHTOKEN:?set NGROK_AUTHTOKEN env var first}"
Optional ngrok exposure:
ngrok http --domain=batnlp.ngrok.app 5656
π¦ Python Packages
pip install --upgrade pip
pip install jupyter pandas openpyxl
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126
cd v2_HE_Universe
Install all dependencies
pip install -r requirements.txt
pip install faster-whisper soundfile pydub speechbrain \
huggingface_hub==0.16.4 einops lightning lightning_utilities torchmetrics primePy
pip install mtcnn opentelemetry-exporter-otlp
pip install --no-deps facenet-pytorch scikit-learn rich pandas matplotlib tensorboard pyannote.audio threadpoolctl torchcodec safetensors pyannote.metrics colorlog pyannote.pipeline optuna pyannote.core pyannote.database \
sortedcontainers torch-audiomentations julius torch-pitch-shift \
opentelemetry-instrumentation opentelemetry-api opentelemetry-distro
pip install asteroid-filterbanks python-dotenv
pip install --no-deps noisereduce lazy_loader audioread soxr numba llvmlite
opentelemetry-bootstrap -a install
# pip install -U pyannote.audio
python get_nltk_data.py
π Audio Backend
sed -i 's/available_backends = .*/available_backends = ["sox_io", "soundfile"]/' \
$CONDA_PREFIX/lib/python3.11/site-packages/speechbrain/utils/torch_audio_backend.py
π§ Patch face_recognition_models
python - <<'EOF'
import pathlib
import importlib.util
file_path = pathlib.Path(importlib.util.find_spec("face_recognition_models").origin)
content = file_path.read_text()
if "pkg_resources" in content:
print("π©Ή Applying safe patch for face_recognition_models...")
new_import = (
"import importlib.resources as resources\n"
"def resource_filename(package_or_requirement, resource_name):\n"
" return str(resources.files(package_or_requirement).joinpath(resource_name))\n"
)
patched = []
for line in content.splitlines():
if "pkg_resources" in line and "import" in line:
patched.append(new_import)
else:
patched.append(line)
file_path.write_text("\n".join(patched))
print("β
Safe patch applied successfully.")
else:
print("β
face_recognition_models already safe or patched.")
EOF
π Secrets & Credentials (not tracked in git)
These files are required at runtime but are gitignored β obtain them from the team and place them locally, they are never committed:
tools/.envβELEVENLABS_API_KEY,FERNET_KEY,ENCRYPTED(HF token),GOOGLE_API_KEYtools/service_account.jsonβ GCP service account for Translatemodules/ekyc/utils/cloudVisionAPI.jsonβ GCP service account for VisionNGROK_AUTHTOKENenv var, if exposing via ngrok
π HuggingFace Hub Login & Token
pip install huggingface_hub==0.16.4
huggingface-cli login
π₯ Fix CUDA Library Path for GPU (cublas/cudnn)
faster-whisper (via ctranslate2) dlopen's libcublas.so.12 at runtime. If
torch was installed from a cu13x index, only libcublas.so.13 is present, so
transcription fails with RuntimeError: Library libcublas.so.12 is not found.
Install the matching CUDA 12 cublas package and register it permanently via a conda activation hook (works in every new shell/reboot, not just this one):
pip install nvidia-cublas-cu12
mkdir -p "$CONDA_PREFIX/etc/conda/activate.d" "$CONDA_PREFIX/etc/conda/deactivate.d"
cat > "$CONDA_PREFIX/etc/conda/activate.d/env_vars.sh" <<'HOOK'
#!/bin/bash
export _HE_OLD_LD_LIBRARY_PATH="$LD_LIBRARY_PATH"
_HE_NVIDIA_LIB_DIR="$CONDA_PREFIX/lib/python3.11/site-packages/nvidia"
_HE_CUDA_LIBS=""
for d in cublas/lib cudnn/lib cuda_runtime/lib cuda_nvrtc/lib; do
[ -d "$_HE_NVIDIA_LIB_DIR/$d" ] && _HE_CUDA_LIBS="$_HE_NVIDIA_LIB_DIR/$d:$_HE_CUDA_LIBS"
done
export LD_LIBRARY_PATH="${_HE_CUDA_LIBS}${LD_LIBRARY_PATH}"
unset _HE_NVIDIA_LIB_DIR _HE_CUDA_LIBS
HOOK
cat > "$CONDA_PREFIX/etc/conda/deactivate.d/env_vars.sh" <<'HOOK'
#!/bin/bash
export LD_LIBRARY_PATH="$_HE_OLD_LD_LIBRARY_PATH"
unset _HE_OLD_LD_LIBRARY_PATH
HOOK
conda deactivate && conda activate HE # re-run the hook
python Universe_API.py