# CISM on Linux — 2-core Intel lake playbook Target: 2-core Intel lake-class CPU (old lake: AVX2+FMA, no VNNI/AVX512). Windows MSVC box is the dev machine; this doc is the Linux test checklist. ## 1. Prereqs - Python 3.11–3.13 (`requires-python = ">=3.11,<3.14"` in `pyproject.toml`). Check: `python3 --version`. - C++ toolchain: `cmake` (>=3.21 recommended), plus `g++` or `clang++` with C++20 support. Check: `cmake --version; g++ --version`. - `pip` (>=23 recommended) and `git`. Check: `pip --version`. - Debian/Ubuntu one-liner: ```sh sudo apt update && sudo apt install -y python3 python3-pip cmake g++ git ``` - No MSVC / vswhere / vcvars needed. `torch.compile` needs no vcvars on Linux (`_ensure_msvc` is a no-op there). ## 2. Build ```sh git clone cd CISM pip install -e .[test] ``` - `.[test]` pulls `pytest/httpx/openai/fastapi/uvicorn` for the harness. - Reference baselines (HF eager/compile) additionally need `pip install -e .[reference]` (torch). Skip torch for CISM-only numbers. - Native kernels build via scikit-build-core (Release). No manual `-mavx2` flags: per-file codegen in `CMakeLists.txt` handles dispatch. ## 3. Quick checks ```sh python -m pytest tests/test_autobench.py tests/test_autobench_rigor.py -q # expect: green (platform helpers are pure Python, no native build needed) python -c "from cism._native import kernel_variant; print(kernel_variant())" # expect on old lake: avx2 (NOT vnni/avx512 — old lakes have no VNNI) ``` - `kernel_variant() == "avx2"` is the pass condition on this box. A `scalar` result means AVX2 dispatch failed — report the cmake log. - `python -c "from cism import Engine; ..."` smoke-loads a small checkpoint once the native build exists. ## 4. Bench commands (2-core box) int8 first (best speed and quality without VNNI); threads 1–2 only; `--reps` defaults to 3 in autobench, `--runs 5` is the CLI legacy name: ```sh cism bench --precision int8 --threads 1 --max-tokens 64 --reps 3 cism bench --precision int8 --threads 2 --max-tokens 64 --reps 3 # then, only if asked: fp32 and/or hybrid-int4 at threads 1-2 for the delta cism bench --precision fp32 --threads 2 --max-tokens 64 --reps 3 ``` - Do NOT sweep 4/6/12 threads here: 2 physical cores, oversubscription convoys. Thread sweep `(1, 2, 4)` in autobench is for bigger boxes. - Keep `--max-tokens 64` on laptops (short = less heat, same ranking). - Prefer `local-files-only` / warm HF cache on metered links. ## 5. Notes - `lock_pages` needs `RLIMIT_MEMLOCK`: check `ulimit -l` (KiB). If it returns 64 or fails, either raise it (`ulimit -l unlimited`, may need `/etc/security/limits.conf`) or skip `--lock-pages` — decode still runs, pages just stay pageable. - Thermals on laptops: a >15% max–min spread across reps flags `thermal_throttled` in the report. Cooldown between thread counts, charger connected, no parallel builds during measurement. - No `wmic` on Linux by design. `_cpu_snapshot()` uses psutil when present (clock/temp) and returns `{}` without psutil — telemetry never fails the run. Missing `/sys` cache files likewise yield `l3_cache_mb: null`, never a crash. ## 6. What numbers to report back Per (model, precision, threads) row, paste from the autobench report: - `model`, `threads`, `K` (spec_k; 0 = off), `precision`/`act`. - `tokens_per_second` (median) **plus** `rep_rates` (all reps), `tokens_per_second_min/max/std`, `reps`, `thermal_throttled`, `thermal_drop_percent`. - `perplexity` (+ `ppl_delta_percent` vs HF FP32 where present). - `cpu_snapshot` (`cpu_mhz` / `cpu_temp_c`, or `{}` if unavailable). - `kernel_variant`, `l3_cache_mb`/`fits_l3` when present. Copy-paste template: ```text CISM rev: CPU: cores: 2 kernel_variant: Model: precision=int8 K=<0|..> reps=3 max-tokens=64 1T: rep_rates=[..] thermal_throttled= PPL=<..> 2T: rep_rates=[..] thermal_throttled= PPL=<..> cpu_snapshot: <{..} or {}> ulimit -l: lock_pages: Notes: ```