You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Prefered System :: NVIDIA RTX A4000 | NVIDIA RTX 3090 Ti

Avoid Servers

  • RTX 5060 Ti
  • Tesla P100

🐳 Docker (recommended β€” run this project on any machine)

docker-compose.yml runs standalone on any machine with plain Docker Desktop (CPU-only β€” torch/ctranslate2/YOLO all fall back automatically, just slower). GPU acceleration is opt-in via docker-compose.gpu.yml, which requires the NVIDIA Container Toolkit on the host and a driver supporting CUDA 13 forward compat (>=580.x). Everything below is a one-time host setup β€” the image itself is fully self-contained and pinned.

# 1. Pull the actual model weights (image does NOT bake these in β€” 5GB, LFS)
git lfs pull

# 2. Provide secrets (never baked into the image β€” see .env.example)
cp .env.example tools/.env   # then fill in real values
# also place: tools/service_account.json, modules/ekyc/utils/cloudVisionAPI.json

# 3. Build + run β€” CPU-only, any machine:
docker compose up --build -d

# ...or with GPU, on a host with the NVIDIA Container Toolkit installed:
docker compose -f docker-compose.yml -f docker-compose.gpu.yml up --build -d

# 4. Check it's alive
curl http://localhost:5656/docs
docker compose logs -f

To stop: docker compose down (add -v to also wipe the model-download caches in speechbrain_cache/hf_cache, forcing a re-download next start).

Verifying the Docker setup

bash scripts/lint-docker.sh        # hadolint + compose config, seconds
bash scripts/docker-smoke-test.sh  # full build + start + wait-for-healthy

docker-smoke-test.sh needs the same prerequisites as step 1-2 above (real AI_Models/ weights + credential files) since the app won't start without them β€” that's also why CI (.github/workflows/docker-ci.yml) only lints and builds, not runs, the image.

Running with plain docker run (no compose, e.g. on a machine that only pulled the image)

Compose's volumes/env/GPU flags translated 1:1:

docker run -d \
  --name he-universe \
  --gpus all \
  -p 5656:5656 \
  -e PUID=$(id -u) -e PGID=$(id -g) \
  -v "$(pwd)/AI_Models:/app/AI_Models:ro" \
  -v speechbrain_cache:/app/AI_Models/SpeechAnalyzer_Models/pretrained_models \
  -v hf_cache:/home/appuser/.cache/huggingface \
  -v "$(pwd)/logs:/app/logs" \
  -v "$(pwd)/tools/.env:/app/tools/.env:ro" \
  -v "$(pwd)/tools/service_account.json:/app/tools/service_account.json:ro" \
  -v "$(pwd)/modules/ekyc/utils/cloudVisionAPI.json:/app/modules/ekyc/utils/cloudVisionAPI.json:ro" \
  --restart unless-stopped \
  hawkeyes-universe:latest

(swap the final image name for yourdockerhubuser/hawkeyes-universe:latest if pulling from Docker Hub instead of building locally)

Publishing to Docker Hub

The image is already built to be push-safe: AI_Models/ (proprietary weights), all credential files, .git, and the Dockerfile/compose files themselves are excluded from the build context (.dockerignore) β€” only application code and pinned dependencies get baked in. Application source code IS baked in though (COPY . .), so unless this is meant to be public, push to a private Docker Hub repository.

export IMAGE_NAME=yourdockerhubuser/hawkeyes-universe
export IMAGE_TAG=1.0.0   # or "latest" β€” pick something you can track releases by

docker login

docker compose build
docker compose push

(docker compose push works because image: in docker-compose.yml is already registry-qualified via $IMAGE_NAME/$IMAGE_TAG.)

To pull and run it on another machine afterward, that machine needs docker login too (for a private repo), then the same docker run/docker compose up commands above with IMAGE_NAME/IMAGE_TAG set to match, plus its own copy of AI_Models/ (git-lfs) and the secret files β€” those never travel with the image.


⚑ Quick Start + 🐍 Miniconda & Conda Environment + πŸ›  Setup NGROK

(manual/bare-metal setup β€” skip this if you're using Docker above)

  dev

sudo apt update && sudo apt upgrade -y
sudo apt install -y iproute2 libgl1 nano wget unzip nvtop git git-lfs build-essential cmake \
libopenblas-dev liblapack-dev libx11-dev libgtk-3-dev libglib2.0-0

git config --global credential.helper store
git clone -b dev https://huggingface.co/HawkEyesAI/v2_HE_Universe

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
chmod +x Miniconda3-latest-Linux-x86_64.sh
./Miniconda3-latest-Linux-x86_64.sh -b -p $HOME/miniconda3
export PATH="$HOME/miniconda3/bin:$PATH"
conda init
source ~/.bashrc
conda create --name HE python=3.11 -y
conda activate HE

curl -s https://ngrok-agent.s3.amazonaws.com/ngrok.asc | sudo tee /etc/apt/trusted.gpg.d/ngrok.asc >/dev/null
echo "deb https://ngrok-agent.s3.amazonaws.com buster main" | sudo tee /etc/apt/sources.list.d/ngrok.list
sudo apt update
sudo apt install -y ngrok
ngrok config add-authtoken "${NGROK_AUTHTOKEN:?set NGROK_AUTHTOKEN env var first}"

Optional ngrok exposure:

ngrok http --domain=batnlp.ngrok.app 5656

πŸ“¦ Python Packages

pip install --upgrade pip
pip install jupyter pandas openpyxl

pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126

cd v2_HE_Universe

Install all dependencies

pip install -r requirements.txt

pip install faster-whisper soundfile pydub speechbrain \
huggingface_hub==0.16.4 einops lightning lightning_utilities torchmetrics primePy

pip install mtcnn opentelemetry-exporter-otlp

pip install --no-deps facenet-pytorch scikit-learn rich pandas matplotlib tensorboard pyannote.audio threadpoolctl torchcodec safetensors pyannote.metrics colorlog pyannote.pipeline optuna pyannote.core pyannote.database \
sortedcontainers torch-audiomentations julius torch-pitch-shift \
opentelemetry-instrumentation opentelemetry-api opentelemetry-distro

pip install asteroid-filterbanks python-dotenv


pip install --no-deps noisereduce lazy_loader audioread soxr numba llvmlite


opentelemetry-bootstrap -a install


# pip install -U pyannote.audio


python get_nltk_data.py

πŸ›  Audio Backend

sed -i 's/available_backends = .*/available_backends = ["sox_io", "soundfile"]/' \
$CONDA_PREFIX/lib/python3.11/site-packages/speechbrain/utils/torch_audio_backend.py

πŸ”§ Patch face_recognition_models

python - <<'EOF'
import pathlib

import importlib.util
file_path = pathlib.Path(importlib.util.find_spec("face_recognition_models").origin)
content = file_path.read_text()

if "pkg_resources" in content:
    print("🩹 Applying safe patch for face_recognition_models...")
    new_import = (
        "import importlib.resources as resources\n"
        "def resource_filename(package_or_requirement, resource_name):\n"
        "    return str(resources.files(package_or_requirement).joinpath(resource_name))\n"
    )
    patched = []
    for line in content.splitlines():
        if "pkg_resources" in line and "import" in line:
            patched.append(new_import)
        else:
            patched.append(line)
    file_path.write_text("\n".join(patched))
    print("βœ… Safe patch applied successfully.")
else:
    print("βœ… face_recognition_models already safe or patched.")
EOF

πŸ”’ Secrets & Credentials (not tracked in git)

These files are required at runtime but are gitignored β€” obtain them from the team and place them locally, they are never committed:

  • tools/.env β€” ELEVENLABS_API_KEY, FERNET_KEY, ENCRYPTED (HF token), GOOGLE_API_KEY
  • tools/service_account.json β€” GCP service account for Translate
  • modules/ekyc/utils/cloudVisionAPI.json β€” GCP service account for Vision
  • NGROK_AUTHTOKEN env var, if exposing via ngrok

πŸ”‘ HuggingFace Hub Login & Token

pip install huggingface_hub==0.16.4
huggingface-cli login

πŸ–₯ Fix CUDA Library Path for GPU (cublas/cudnn)

faster-whisper (via ctranslate2) dlopen's libcublas.so.12 at runtime. If torch was installed from a cu13x index, only libcublas.so.13 is present, so transcription fails with RuntimeError: Library libcublas.so.12 is not found.

Install the matching CUDA 12 cublas package and register it permanently via a conda activation hook (works in every new shell/reboot, not just this one):

pip install nvidia-cublas-cu12

mkdir -p "$CONDA_PREFIX/etc/conda/activate.d" "$CONDA_PREFIX/etc/conda/deactivate.d"
cat > "$CONDA_PREFIX/etc/conda/activate.d/env_vars.sh" <<'HOOK'
#!/bin/bash
export _HE_OLD_LD_LIBRARY_PATH="$LD_LIBRARY_PATH"
_HE_NVIDIA_LIB_DIR="$CONDA_PREFIX/lib/python3.11/site-packages/nvidia"
_HE_CUDA_LIBS=""
for d in cublas/lib cudnn/lib cuda_runtime/lib cuda_nvrtc/lib; do
    [ -d "$_HE_NVIDIA_LIB_DIR/$d" ] && _HE_CUDA_LIBS="$_HE_NVIDIA_LIB_DIR/$d:$_HE_CUDA_LIBS"
done
export LD_LIBRARY_PATH="${_HE_CUDA_LIBS}${LD_LIBRARY_PATH}"
unset _HE_NVIDIA_LIB_DIR _HE_CUDA_LIBS
HOOK
cat > "$CONDA_PREFIX/etc/conda/deactivate.d/env_vars.sh" <<'HOOK'
#!/bin/bash
export LD_LIBRARY_PATH="$_HE_OLD_LD_LIBRARY_PATH"
unset _HE_OLD_LD_LIBRARY_PATH
HOOK

conda deactivate && conda activate HE  # re-run the hook
python Universe_API.py
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support