omniAI / docs /model-interface.md
hasimjaneef's picture
Publish validated findings before turn end and trace pipeline latency
79e11ca verified
|
Raw History Blame Contribute Delete
20.1 kB

MiniCPM-o 4.5 native interface and implementation boundary

Reviewed on 2026-09-08; text-only startup rechecked on 2026-09-09; live-output control and block delivery rechecked on 2026-09-10. This application uses only openbmb/MiniCPM-o-4_5 for audiovisual understanding and portfolio conclusions. Python validates allocations, records source evidence, combines public output, suppresses duplicate publications, and assembles the final brief. There is no second model, separate portfolio role, translation request, question-generation request, JSON-repair request, or final-summary request.

Immutable sources

The official model card was checked first, followed by its inference examples and actual implementation. The seven reference files and five model source files were downloaded again and their SHA256 hashes verified locally. No downloaded inference code was executed in that check.

Source Pinned revision and relevant files
Required official model 503e754207c94da6bb26850b4469f367c9ea3582; model weights and remote Python use the same snapshot.
Native model implementation MiniCPMO.as_duplex, MiniCPMODuplex.prepare, streaming_prefill, streaming_generate, set_session_stop.
Audio processor StreamingMelProcessorExact, chunk alignment and sliding waveform buffer.
Decoder and window implementation StreamDecoder.decode, unit registration, protected system prefix, basic KV window.
Official main repository f0866559fae0305bc7cacfb6a950640a927f6984; README inference examples. The live repository URL now redirects to MiniCPM-V; this does not change the explicitly pinned MiniCPM-o model.
Official Demo 50b0865c819c2f0ca24ec7994e05044e5f39d451; core/processors/unified.py, MiniCPMO45/modeling_minicpmo_unified.py, and English duplex architecture/video-protocol docs.

model-lock.json records immutable revisions and all reviewed fingerprints. python scripts/verify_model_sources.py fetches only these small public source/reference files, validates their hashes, and checks important AST contracts. It does not fetch weights or read credentials.

Selected native lifecycle

from gpu_service.native import prepare_text_context, text_only_duplex, begin_public_response
from threading import Event

model = AutoModel.from_pretrained(
    pinned_local_snapshot, trust_remote_code=True, local_files_only=True,
    attn_implementation="sdpa", torch_dtype=torch.bfloat16,
    init_vision=True, init_audio=True, init_tts=False,
).eval().cuda()
duplex = text_only_duplex(model,
    generate_audio=False, chunk_ms=1000, first_chunk_ms=1035,
    sample_rate=16000, sliding_window_mode="basic",
    basic_window_high_tokens=8000, basic_window_low_tokens=6000,
)
model.reset_session()
prepare_text_context(duplex, portfolio_and_objective_prompt, Event())  # once per epoch
duplex.streaming_prefill(audio_waveform=pcm, frame_list=rgb_frames,
                         max_slice_nums=1, batch_vision_feed=True)
# On the first real unit; NativeDuplex handles later starts and active responses.
begin_public_response(duplex)
duplex.streaming_generate(max_new_speak_tokens_per_chunk=20,
                          decode_mode="sampling", listen_prob_scale=1.0,
                          listen_top_k=None)

The model is loaded once into a persistent CUDA process, with one active authorized session. Each audiovisual second uses the same decoder, audio encoder state, and protected session context. There is no standard image-chat endpoint in this path.

text_only_duplex calls the same official model.as_duplex after temporarily replacing only that model instance's init_tts method with a no-op. It restores the method in finally, including on constructor failure. This compatibility shim is necessary because the pinned constructor calls init_tts unconditionally even when generate_audio=False; the upstream interface has no separate skip-renderer flag. Source fingerprints are verified before loading. Audio/vision encoders, the language decoder, audiovisual prefill, generation and session controls remain the official implementation.

The application-owned prepare_text_context follows the pinned prepare text-only/basic-window path: clear native stop flags, reset streaming state and decoder, initialize the streaming audio processor, feed the exact system-prefix and suffix tokens, and register the protected context boundary. It batches at most 256 prompt tokens into each supported StreamDecoder.embed_tokens / feed call instead of upstream's one full decoder forward per prompt token. Cancellation is checked between batches. No answer is generated and no model role is changed. Reference audio and the alternative context-preserving window are excluded by this adapter. A CPU fixture executed the exact pinned method and the new path with inert substitutes, verifying equal token order, reset order and protected boundary. This does not measure CUDA latency or establish numerical equivalence on hardware.

Audiovisual prefill uses the official batch_vision_feed=True option to combine a frame's image markers and embeddings into one decoder feed. Upstream documents possible small floating-point differences from sequential feeds. Input frames, audio, source cadence and native unit finalization are unchanged.

The adapter reads window metadata directly from decoder.get_cache_length() and len(decoder._unit_history). The pinned optional get_window_stats() helper references _system_prompt_template at utils.py:2000, but neither its constructor nor reset defines that attribute. Calling that helper before the first prefill caused the reported HTTP 502. Reading the same two authoritative values avoids the defective diagnostic branch without inventing statistics, changing the cache, modifying upstream files or disabling context-window enforcement. The exact pinned constructor/reset/statistics methods reproduced this error locally with inert model/cache substitutes; the revised reads succeeded before input, with retained units, and after reset.

Do not mix the two official implementations. The selected Hugging Face streaming_generate closes </unit>, registers the unit, and applies the window before returning, including listening units. Its signature has no force_listen_override or separate finalize_unit requirement. The pinned Demo has those different APIs and requires finalization. Its documented 300-second gateway limit is a Demo policy; OmnAI does not use that gateway or inherit the limit.

Audio, video and time

Item Verified input and application behavior
Audio transport Base64 of little-endian float32 PCM; mono, normalized finite samples in [-1, 1]; exactly 16,000 Hz. No WAV header in a chunk.
Native audio array One-dimensional NumPy float32 array; normally 16,000 samples per one-second unit. Both audio and frames are passed to the same streaming_prefill call, which selects native OMNI mode.
Video transport Bounded JPEGs, decoded into PIL RGB images. App policy caps frames at 1920×1080 and 512,000 encoded bytes each. The browser normally captures at most one sampled frame per media second. Missing frames remain an explicit coverage limitation.
Vision processing max_slice_nums=2 reaches the active prefill call. This is a maximum, not a demand for two crops. The pinned processor returns no extra grid when image area is at most 448×448; the unchanged browser longest-edge cap is 448, so browser frames still receive only their overview. Larger direct-service inputs can use extra slices. Raising resolution is outside this patch. Native image processing handles dimensions/aspect ratio. General model-card high-FPS/high-resolution claims are not a measured cadence or OCR guarantee for this configuration.
Cadence Native prefill and listen/speak decision are called once for each second of source audio. A local video may be paused to pace inference. This is continuous sampled audiovisual observation, not every-frame analysis.
First-unit alignment Stock native code prepends 560 zero samples to 16,000 input samples, then consumes a 16,480-sample core. It retains 80 samples (5 ms) for the following unit. The adapter tracks synthetic padding separately from source samples.
Final partial unit The application sends every real sample with its exact count; the adapter right-pads to one native second and records the padding as synthetic. Tiny final fragments remain accepted.
Natural drain If any real samples remain in the native buffer, a synthetic zero-audio unit drains them. Up to eight synthetic units may finish the current utterance; they add no source time and carry no new visual evidence. The report states whether audio is fully consumed and whether the utterance remained incomplete.
Timestamps Frame PTS and audio source start/end are application metadata. Source milliseconds retain their measured precision. Native current_time is a unit counter, not a media timestamp or real-world event time. Capture, backend receipt, inference start/end, and emission clocks remain separate.

The public text is an audiovisual finding or Audio evidence, never a promised verbatim transcript. Native output does not establish word-aligned speech timestamps. Conclusions receive the conservative union of the source intervals and evidence IDs available across their generation, rather than being attributed solely to the latest frame. These references identify available context; the native model does not provide exact word-to-frame attribution.

Context, generation cost and bounded state

POST /sessions accepts session_id, epoch, objective, holdings, output_max_tokens_per_unit (8–64, default 20), and live_update_interval_seconds (1–30 source seconds, default 5). Each holding contains id, name, ticker, exchange, currency, asset_type, and percentage weight. Code validates weights total 100% and identities are unique. The authenticated service can accept an empty portfolio for a basic audiovisual smoke test; the application requires holdings.

Holdings and objective are supplied in one compact system prompt per source epoch. prepare resets native caches, so seek/reconnect epochs prepare this context again. A repeated identical start is idempotent; changing context in-place requires a newer epoch. The protected prompt is limited to 4,000 actual tokenizer tokens so it cannot consume the entire configured KV window. Context that exceeds this limit must be shortened; it is never silently truncated.

The native decoder preserves the system prompt and removes old audiovisual units when its 8,000-token high watermark is crossed, targeting 6,000 tokens. This does not mean it retains all ten minutes simultaneously. The audio Mel processor enables its upstream 30-second sliding buffer with 10-second drops; the audio encoder resets its own KV when it reaches its positional capacity. Application evidence and concise completed findings persist separately in bounded buffers for the final brief. Unused native debug/TTS tensor history is cleared after each unit; the recent 512-token repetition window remains intact.

MiniCPM retains its own native listen/speak decision on every unit. The previous application override of pending_logits was removed at the user's request to preserve native silence. listen_top_k=None and listen_prob_scale=1.0 remain unchanged. The legacy live_update_interval_seconds request field is accepted for compatibility but does not force output or impose a publication delay. The portfolio/objective prompt is unchanged by this patch.

Complete labeled portfolio findings can enter the public ledger before end_of_turn: the existing label grammar must identify supplied holdings, evidence, implications, uncertainty and verification actions, and the last field must have a closed line/block and sentence punctuation. Incomplete or ambiguous tails remain buffered. Evidence-only/unstructured output retains the existing turn-end fallback. Public-output safety checks and the existing material-term-aware deduplication still apply. A continuing turn updates the same evidence/assessment record; its complete text is retained at turn-end, without duplicate final-brief entries. Arbitrary partial text is no longer displayed through live_observation.

Authenticated revision-based long polling already wakes immediately on changes. No publication timer, extra summary model, speech synthesis, or additional model request is introduced. Natural completion retains its bounded tail; it is reported separately from ordinary input processing. The output cap is unchanged: 20 native tokens per unit by default, configurable 8–64. Lowering that cap could spread one finding across more source units.

The GPU service decodes PCM and JPEG once on a CPU executor, validates before model mutation, then submits the decoded arrays/images to the existing single native worker. It rejects inputs invalidated by Stop/seek during decoding. The application backend still performs its own bounded admission validation; this patch removes duplicate decoding within the GPU service, not validation at separate trust boundaries.

Process-local monotonic timing records distinguish CPU preparation, worker wait, native prefill, generation, application queue, validation, browser send/receipt and render opportunity. Existing native cost_* fields are included as upstream wall-clock diagnostics, not CUDA-event measurements. No CUDA synchronization is added. Effective component devices/dtypes/attention classes and window/output settings are reported at preparation. See latency findings and validation procedure.

The native parameter named listen_prob_scale actually multiplies a raw logit in the pinned source. A value above 1 can decrease listening preference when that logit is negative; OmnAI therefore keeps it at 1.0. No undocumented forced-silence tokens are inserted. Every accepted unit still completes the native generation/finalization step; simply skipping it would corrupt the unit sequence.

The configured token limit is per native unit, including native control-token handling, not a promise of that many visible words. A concise finding can span multiple seconds. Removing the portfolio model eliminates its separate inference calls; audiovisual encoders, KV updates, and native listen decisions still consume GPU compute.

Stop, failure and completion

Stop sets upstream set_session_stop() and the application cancellation event immediately. Native prefill/generation check stop events at call boundaries; the pinned text decoding loop is bounded but does not promise hard mid-kernel interruption. The service discards late outputs, rejects additional chunks/finish, and queues cache reset on the same single worker after in-flight work returns. It never mutates a running CUDA decoder from a second thread. Hardware stalls remain outside an application cancellation guarantee.

Seeking creates a strictly newer epoch. Ordered sequences, exact retry-payload hashes, generation IDs and source continuity checks reject stale, out-of-order or conflicting results. At most one call enters the GPU worker; concurrent requests receive bounded backpressure. Sixteen recent completed transport results support idempotent retry without a second inference call. Timeout makes the epoch uncertain: the application fails the session, cancels pending work, asks for native reset and opens a partial brief. It does not automatically replay the source, create another epoch or resend the portfolio prompt. A new Start is explicit. Pending counts include the currently executing request, so playback pacing can engage while the first request is still running.

Natural completion flushes real audio plus the bounded native utterance tail. It makes zero final-summary requests. Python assembles the research brief from already completed public findings. Manual Stop makes no tail or summary requests and presents a partial brief. Stopping analysis does not pause dedicated GPU uptime billing.

Runtime pins and verification status

The selected target is one NVIDIA A100 with 80 GB VRAM. The pinned installer is pip 25.3, bootstrapped from PyPI. The CUDA-index step installs only the two exact wheels with --no-deps --only-binary=:all:; dependencies are installed afterward from PyPI and checked, including the installed CUDA package versions. The pinned runtime is Python 3.10.18, Torch/Torchaudio 2.8.0 from the CUDA 12.8 wheel index, Transformers 4.51.0, minicpmo-utils[all] 1.0.6, NumPy 1.26.4, Pillow 10.4.0 and PyTorch SDPA. All remaining package versions are pinned in requirements-gpu.txt. Linux x86_64/Python 3.10 dependency resolution was repeated successfully on 2026-09-08. This is metadata compatibility, not a CUDA execution test.

The model card specifies Transformers 4.51.0 and a Torch version no later than 2.8.0. The utility package forces Librosa 0.9.0 and Pillow 10.4.0; Setuptools 80.9.0 retains pkg_resources needed by that older Librosa. The separate Demo's newer broad requirements are not mixed into this selected Hugging Face runtime.

generate_audio=False returns before speech-token decoding and Token2Wav synthesis. Together with init_tts=False and the compatibility shim above, this avoids constructing the unused TTS module and Token2Wav renderer. Snapshot downloads exclude assets/**; the same four model checkpoint shards are still required. No speech is synthesized. The native listen branch may construct a zero-filled silence array, which the application ignores. Dependencies remain pinned to the previously built combination; this change does not replace the audio encoder or introduce another model.

HF_TOKEN is read server-side at runtime: the GPU process uses it for model download, and an optional application-side token can authenticate a private Hugging Face Space gateway. The snapshot is pinned and local subsequent loads are used. INFERENCE_SERVICE_TOKEN independently protects the service; URLs and tokens are handled server-side by the application. Neither credential is built into Docker layers or frontend assets.

Verified locally: official-source fingerprints and AST contracts; pinned dependency resolution; CPU native-shaped PCM/JPEG request, context, cadence, tail, Stop, ordering and epoch tests. A separate fixture executed the exact pinned constructor and prepare methods with inert decoder/tensor substitutes, confirming the speech-renderer call is skipped and the streaming audio processor and protected prompt are initialized. This is not model inference. The simulated ten-minute test validates 600 unique audio chunks plus all final samples and lifecycle budgets.

Observed in the supplied GPU runtime log starting 2026-09-10 02:03:18 UTC: all four checkpoint shards loaded and Uvicorn startup completed. The first audiovisual request then failed at _feed → _stats → utils.py:2000:get_window_stats, identifying the missing _system_prompt_template attribute described above. This supersedes the earlier unconfirmed-error diagnosis. The exact crash has now been reproduced and corrected locally. The updated adapter still needs upload and a short authorized audiovisual GPU run to establish successful real inference; active-session memory use, 1× realtime throughput and English/portfolio output quality remain unverified. No ten-minute GPU performance claim is made.