test1111111 / docs /LINUX.md
spitfire4794's picture
CISM remote autobench: full source + fleet runner, serve results on 7860
28a1a01
|
Raw History Blame Contribute Delete
4.18 kB

CISM on Linux — 2-core Intel lake playbook

Target: 2-core Intel lake-class CPU (old lake: AVX2+FMA, no VNNI/AVX512). Windows MSVC box is the dev machine; this doc is the Linux test checklist.

1. Prereqs

  • Python 3.11–3.13 (requires-python = ">=3.11,<3.14" in pyproject.toml). Check: python3 --version.
  • C++ toolchain: cmake (>=3.21 recommended), plus g++ or clang++ with C++20 support. Check: cmake --version; g++ --version.
  • pip (>=23 recommended) and git. Check: pip --version.
  • Debian/Ubuntu one-liner:
sudo apt update && sudo apt install -y python3 python3-pip cmake g++ git
  • No MSVC / vswhere / vcvars needed. torch.compile needs no vcvars on Linux (_ensure_msvc is a no-op there).

2. Build

git clone <cism-repo-url>
cd CISM
pip install -e .[test]
  • .[test] pulls pytest/httpx/openai/fastapi/uvicorn for the harness.
  • Reference baselines (HF eager/compile) additionally need pip install -e .[reference] (torch). Skip torch for CISM-only numbers.
  • Native kernels build via scikit-build-core (Release). No manual -mavx2 flags: per-file codegen in CMakeLists.txt handles dispatch.

3. Quick checks

python -m pytest tests/test_autobench.py tests/test_autobench_rigor.py -q
# expect: green (platform helpers are pure Python, no native build needed)

python -c "from cism._native import kernel_variant; print(kernel_variant())"
# expect on old lake: avx2 (NOT vnni/avx512 — old lakes have no VNNI)
  • kernel_variant() == "avx2" is the pass condition on this box. A scalar result means AVX2 dispatch failed — report the cmake log.
  • python -c "from cism import Engine; ..." smoke-loads a small checkpoint once the native build exists.

4. Bench commands (2-core box)

int8 first (best speed and quality without VNNI); threads 1–2 only; --reps defaults to 3 in autobench, --runs 5 is the CLI legacy name:

cism bench <model-id> --precision int8 --threads 1 --max-tokens 64 --reps 3
cism bench <model-id> --precision int8 --threads 2 --max-tokens 64 --reps 3
# then, only if asked: fp32 and/or hybrid-int4 at threads 1-2 for the delta
cism bench <model-id> --precision fp32 --threads 2 --max-tokens 64 --reps 3
  • Do NOT sweep 4/6/12 threads here: 2 physical cores, oversubscription convoys. Thread sweep (1, 2, 4) in autobench is for bigger boxes.
  • Keep --max-tokens 64 on laptops (short = less heat, same ranking).
  • Prefer local-files-only / warm HF cache on metered links.

5. Notes

  • lock_pages needs RLIMIT_MEMLOCK: check ulimit -l (KiB). If it returns 64 or fails, either raise it (ulimit -l unlimited, may need /etc/security/limits.conf) or skip --lock-pages — decode still runs, pages just stay pageable.
  • Thermals on laptops: a >15% max–min spread across reps flags thermal_throttled in the report. Cooldown between thread counts, charger connected, no parallel builds during measurement.
  • No wmic on Linux by design. _cpu_snapshot() uses psutil when present (clock/temp) and returns {} without psutil — telemetry never fails the run. Missing /sys cache files likewise yield l3_cache_mb: null, never a crash.

6. What numbers to report back

Per (model, precision, threads) row, paste from the autobench report:

  • model, threads, K (spec_k; 0 = off), precision/act.
  • tokens_per_second (median) plus rep_rates (all reps), tokens_per_second_min/max/std, reps, thermal_throttled, thermal_drop_percent.
  • perplexity (+ ppl_delta_percent vs HF FP32 where present).
  • cpu_snapshot (cpu_mhz / cpu_temp_c, or {} if unavailable).
  • kernel_variant, l3_cache_mb/fits_l3 when present.

Copy-paste template:

CISM rev: <git rev-parse --short HEAD>
CPU: <lscpu model name>  cores: 2  kernel_variant: <avx2|other>
Model: <id>  precision=int8  K=<0|..>  reps=3  max-tokens=64
1T: <tok/s>  rep_rates=[..]  thermal_throttled=<bool>  PPL=<..>
2T: <tok/s>  rep_rates=[..]  thermal_throttled=<bool>  PPL=<..>
cpu_snapshot: <{..} or {}>
ulimit -l: <value>  lock_pages: <on|off|failed-soft>
Notes: <heat, charger, background load>