Spaces:
Sleeping
Sleeping
|
Download docs/LINUX.md from spitfire4794/test1111111: direct link, hf CLI and curl.
- Browser
- Download file 4.18 kB
-
https://huggingface.co/spaces/spitfire4794/test1111111/resolve/main/docs/LINUX.md
- Command line
-
hf download hf://spaces/spitfire4794/test1111111/docs/LINUX.md
-
curl -L -o LINUX.md https://huggingface.co/spaces/spitfire4794/test1111111/resolve/main/docs/LINUX.md
4.18 kB
CISM on Linux — 2-core Intel lake playbook
Target: 2-core Intel lake-class CPU (old lake: AVX2+FMA, no VNNI/AVX512). Windows MSVC box is the dev machine; this doc is the Linux test checklist.
1. Prereqs
- Python 3.11–3.13 (
requires-python = ">=3.11,<3.14"inpyproject.toml). Check:python3 --version. - C++ toolchain:
cmake(>=3.21 recommended), plusg++orclang++with C++20 support. Check:cmake --version; g++ --version. pip(>=23 recommended) andgit. Check:pip --version.- Debian/Ubuntu one-liner:
sudo apt update && sudo apt install -y python3 python3-pip cmake g++ git
- No MSVC / vswhere / vcvars needed.
torch.compileneeds no vcvars on Linux (_ensure_msvcis a no-op there).
2. Build
git clone <cism-repo-url>
cd CISM
pip install -e .[test]
.[test]pullspytest/httpx/openai/fastapi/uvicornfor the harness.- Reference baselines (HF eager/compile) additionally need
pip install -e .[reference](torch). Skip torch for CISM-only numbers. - Native kernels build via scikit-build-core (Release). No manual
-mavx2flags: per-file codegen inCMakeLists.txthandles dispatch.
3. Quick checks
python -m pytest tests/test_autobench.py tests/test_autobench_rigor.py -q
# expect: green (platform helpers are pure Python, no native build needed)
python -c "from cism._native import kernel_variant; print(kernel_variant())"
# expect on old lake: avx2 (NOT vnni/avx512 — old lakes have no VNNI)
kernel_variant() == "avx2"is the pass condition on this box. Ascalarresult means AVX2 dispatch failed — report the cmake log.python -c "from cism import Engine; ..."smoke-loads a small checkpoint once the native build exists.
4. Bench commands (2-core box)
int8 first (best speed and quality without VNNI); threads 1–2 only;
--reps defaults to 3 in autobench, --runs 5 is the CLI legacy name:
cism bench <model-id> --precision int8 --threads 1 --max-tokens 64 --reps 3
cism bench <model-id> --precision int8 --threads 2 --max-tokens 64 --reps 3
# then, only if asked: fp32 and/or hybrid-int4 at threads 1-2 for the delta
cism bench <model-id> --precision fp32 --threads 2 --max-tokens 64 --reps 3
- Do NOT sweep 4/6/12 threads here: 2 physical cores, oversubscription
convoys. Thread sweep
(1, 2, 4)in autobench is for bigger boxes. - Keep
--max-tokens 64on laptops (short = less heat, same ranking). - Prefer
local-files-only/ warm HF cache on metered links.
5. Notes
lock_pagesneedsRLIMIT_MEMLOCK: checkulimit -l(KiB). If it returns 64 or fails, either raise it (ulimit -l unlimited, may need/etc/security/limits.conf) or skip--lock-pages— decode still runs, pages just stay pageable.- Thermals on laptops: a >15% max–min spread across reps flags
thermal_throttledin the report. Cooldown between thread counts, charger connected, no parallel builds during measurement. - No
wmicon Linux by design._cpu_snapshot()uses psutil when present (clock/temp) and returns{}without psutil — telemetry never fails the run. Missing/syscache files likewise yieldl3_cache_mb: null, never a crash.
6. What numbers to report back
Per (model, precision, threads) row, paste from the autobench report:
model,threads,K(spec_k; 0 = off),precision/act.tokens_per_second(median) plusrep_rates(all reps),tokens_per_second_min/max/std,reps,thermal_throttled,thermal_drop_percent.perplexity(+ppl_delta_percentvs HF FP32 where present).cpu_snapshot(cpu_mhz/cpu_temp_c, or{}if unavailable).kernel_variant,l3_cache_mb/fits_l3when present.
Copy-paste template:
CISM rev: <git rev-parse --short HEAD>
CPU: <lscpu model name> cores: 2 kernel_variant: <avx2|other>
Model: <id> precision=int8 K=<0|..> reps=3 max-tokens=64
1T: <tok/s> rep_rates=[..] thermal_throttled=<bool> PPL=<..>
2T: <tok/s> rep_rates=[..] thermal_throttled=<bool> PPL=<..>
cpu_snapshot: <{..} or {}>
ulimit -l: <value> lock_pages: <on|off|failed-soft>
Notes: <heat, charger, background load>