Reuse shared features for paired 2D and 3D output

#1

This shares Hecate's feature computation when a caller requests both a 2D ink map and 3D depth scores. We encountered the duplicated execution while using ARGUS to inspect scroll predictions across depth. It retains the checkpoint, existing heads, tiling/blending and single-output branches.

On retained PHerc0139 development-exposed material (16x256x256, Hecate 9.6 micrometers, FP32, batch 1, stride 32, two CPU threads, GTX 1660 Ti), three alternating actual-CLI baseline/candidate pairs matched all 65,536 2D pixels and 1,048,576 3D voxels per pair. Feature calls fell from 98 to 49. Median launch-to-exit time fell from 17.8357 to 11.5913 seconds (35.01% less). Starts were 52-54 C, maximum 68 C, and no cooling pause occurred. Peak process-tree RAM was 1.541 GiB and minimum whole-GPU free memory 4,003 MiB.

Four portable synthetic tests are included; run python -m pytest -q test_hecate_shared_output.py. Separate single-output/CLI checks and the data recipe are documented in the public evidence package.

Evidence and reproduction: https://github.com/Cinder-Covenant/ARGUS/blob/18cbbd3e079cf6cfb6278c8c9afdc7a5b773d37a/docs/september_2026/hecate/PUBLIC_REVIEW_DRAFT.md
Numeric measurements: https://github.com/Cinder-Covenant/ARGUS/blob/18cbbd3e079cf6cfb6278c8c9afdc7a5b773d37a/docs/september_2026/hecate/PUBLIC_MEASURED_RESULT_20260930.json

This measures one small paired-output field on one GPU. Startup, identity checks and the same resource hooks are included; raw resampling and UI transport are excluded, and the OS cache was not cleared. This is not a claim of better ink accuracy, unread-scroll generalization, 2D-only speedup, or validated 2.4-micrometer performance. Hecate's architecture, model and baseline belong to the original authors. LLM tools assisted implementation and verification under operator direction.

darthceltic85 changed pull request status to open

Additional actual-CLI check on a retained, exposed PHerc0139 field (16×512×512, 9.6 µm, FP32, batch 1, stride 32, two CPU threads, GTX 1660 Ti 6 GiB): three alternating baseline/candidate fresh-process pairs produced bitwise-equal decoded output in every pair, including all 262,144 2D pixels and 4,194,304 3D voxels. I decoded and compared the six outputs independently after the harness finished. Feature-network calls fell from 450 to 225 per run, with no checkpoint change. The guarded run recorded minimum available RAM 13.198 GiB, minimum whole-GPU free memory 3786 MiB, maximum GPU temperature 70°C, and all six workers stopped.

Median complete-process times were 163.010 s baseline and 57.328 s candidate, a 64.83% observed difference under this guard. This larger timing has a material thermal confound: baseline processes paused for 55.2, 99.1 and 99.2 s, while candidate processes paused for about 22 s each. I would use the previously posted 16×256×256 pause-free 35.01% comparison as the cleaner timing evidence; the larger check strengthens output parity and feature-call scaling, not a whole-scroll speed or ink-accuracy claim. The public machine-readable result and limitations are at https://github.com/Cinder-Covenant/ARGUS/blob/d0eb41bdef13ffb02081f84abde34bf9ac09a88f/docs/september_2026/hecate/ACTUAL_CLI_512_RESULT_20260930.json .

Reproducibility update: the public ARGUS release now includes the guarded PHerc0139 exposed-control preparation, bounded source-pixel verifier, and strict-load-only check. The source correspondence receipt records exact equality for the selected 21x288x288 crop (1,741,824 pixels) after 12 source chunks totaling 5,505,024 bytes; this does not validate the whole volume or establish a fresh-machine acquisition. The strict-load receipt records a 3.981-second Hecate 9.6 µm model load with inference=false. The preparation, verifier and route instructions are linked from the public Hecate control runbook, with preparation, source verification, strict-load verification, and machine-readable receipts. These additions concern setup and reproducibility of a retained exposed control; they do not show better ink accuracy, unread-scroll generalization or independent adoption. The checkpoint and Hecate provider are unchanged.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment