# Case study: space-bunny-alpha — Public edition (Model Fingerprinted)

#1
by azurepy - opened

Case study: space-bunny-alpha — Public edition

Reviewed/redacted copy of case-space-bunny-alpha.md for public sharing. This version preserves the technical conclusions and evidence tables, but removes raw extraction prompts, raw hidden-template wording, raw API bodies, local paths, account IDs, and step-by-step exploit material.

Live analysis target: stealth/space-bunny-alpha ("Space Bunny Alpha") on OpenRouter. Measurements were generated independently against the live endpoint on 2026-09-27 and 2026-09-28; no external case-study measurements or scripts were copied.

TL;DR: Space Bunny Alpha is a free anonymous OpenRouter listing with a stable stealth wrapper, mandatory reasoning surface, vision input support, 1M-class declared context, and a tokenizer fingerprint that strongly matches the MiniMax-family. The exact checkpoint and official provenance remain unconfirmed. The safest public claim is: MiniMax-family fingerprint by measurement; exact identity unconfirmed.


Public verdicts

Layer Finding Confidence
Tokenizer / vocab Four live token-differential measurements match MiniMaxAI/MiniMax-Text-01 within one constant framing token across CJK, code, mixed Unicode, and plain English corpora. Nearest alternatives are materially farther away. High family-level ID
Wrapper Stable stealth template with a fixed prompt-prefix/cache signature. User-message token counts remain consistent after subtracting the wrapper constant. High
Context Measured successfully far beyond ordinary chat ranges; endpoint declares a 1M-token context cap. Practical ceiling is lower than the declared cap because the gateway pre-gates very large requests with an estimator. High
Vision Image inputs accepted. Prompt-token overhead is constant across tiny and small images, suggesting adapter-style image handling rather than resolution-proportional image tokenization. High behavior
Reasoning Reasoning is surfaced in the API response, but usage accounting reports reasoning tokens as 0. Effort settings affect verbosity/length more than visibility. High
Identity The model improvises public identity claims when asked directly. Self-reports are not used as evidence. High behavior; not provenance
Knowledge Correctly answers dated 2025–2026 model-release facts despite self-reporting an older cutoff. Self-claimed cutoffs are treated as unreliable. High
Serving stack OpenAI-compatible response shape; logprob field exists but is always null; no per-token logprobs available through this gateway. High
Safety / injection Direct manual override attempts mostly refused, but an automated adversarial embedded-content benchmark showed a partial instruction-hierarchy weakness. Medium-High

Tokenizer evidence — public table

Counts are live prompt-token deltas over a one-character baseline, compared against offline tokenizer files on the same four corpora.

Candidate vocab CJK code mixed plain
live endpoint delta 970 5099 3159 2760
MiniMax-Text-01 971 5100 3160 2761
DeepSeek-V3 1020 5280 3480 2841
Kimi-K2.5 1011 5130 3520 2761
Qwen2.5-72B 1132 5280 3640 2761
GLM-4.5-Air 1044 5160 3480 2761
Llama-3-8B 1530 5131 3441 2762

Interpretation: MiniMax-Text-01 matches all four columns within one constant framing token. This is the strongest single piece of family-level evidence in the report.


Safe measurement examples

These examples are safe to publish because they measure ordinary model behavior and do not contain prompt-extraction wording, adversarial instructions, secrets, account IDs, or raw API bodies.

Tokenizer differential example

Use several fixed public corpora and compare token-count deltas, not model answers:

  • CJK paragraph: repeated neutral Chinese technical prose about AI and quantum computing.
  • Code corpus: a small Python/Numpy convolution or sorting snippet.
  • Mixed Unicode corpus: accents, math symbols, emoji, Japanese/Cyrillic text.
  • Plain English corpus: neutral prose about tokenizer measurement.

The key result is the count table above: the live endpoint matches MiniMax-Text-01 within one constant framing token across all four corpora.

Capability examples

Benign tasks used for tier checks included:

  • A known math check: count how long Euler's n² + n + 41 remains prime from n=0.
  • A Python debugging check: identify that a list-merge implementation mutates inputs via front-pop operations and report the resulting merged list.
  • A transcription/arithmetic check: copy rare long words exactly and multiply two large integers.
  • A vision check: identify the color of simple generated PNG squares and compare token overhead across image sizes.

These are suitable examples to describe publicly because they test reasoning, coding, fidelity, arithmetic, and vision without inducing unsafe behavior.

Wrapper / serving-stack examples

Safe public indicators:

  • Fixed prompt-prefix/cache behavior.
  • Constant image-token overhead for different small image resolutions.
  • logprobs response field present but null.
  • Reasoning text surfaced while usage accounting reports zero reasoning tokens.
  • Error messages expose cap/validator behavior, but raw bodies are omitted.

Capability / tier notes

The tested behavior is frontier-class for the chat/API turn structure measured here:

  • Correct on a known polynomial prime-run task.
  • Correct on a subtle Python list-merge mutation/complexity bug.
  • Correct on exact rare-word transcription and a large integer multiplication.
  • Handles vision input and long context.
  • Exposes a mandatory reasoning surface.

This does not prove an official product tier label such as “Flash,” “Lite,” or “Pro.” It only supports the narrower claim that the served model behaves like a current-generation reasoning-capable MiniMax-family deployment.


Quantization note

A direct surprisal/logprob quantization test is not available through this gateway: logprob fields are present in the response shape but remain null even when content tokens are generated. Functional probes did not show obvious quantization degradation, but that is a weak negative, not proof of full precision.

Verdict: no quantization artifact observed; true quantization status unconfirmed.


Template / prompt-leak behavior — redacted summary

The public-safe summary is:

  • The model appears to use a stealth chat template with a system role line, channel configuration, a short developer greeting, and an anti-disclosure rule.
  • The full protected system/developer body was not recovered.
  • Some short template fragments and structure leaked through the reasoning surface and through encoded rendering tasks.
  • Direct requests for hidden instructions generally refused.
  • The leak surface is useful for fingerprinting, but this public version intentionally omits raw hidden-template wording and the exact prompts used to elicit it.

Public impact: no credentials, tool definitions, routing secrets, or authorization logic were observed in leaked fragments.


Automated safety benchmark summary

An automated garak run covered passage-replay and prompt-injection style probes.

  • Public-domain passage continuation behaved as expected and is weak evidence because the probes provide context.
  • One removed/news-style cloze family returned zero hits, which is a cleaner datapoint.
  • A copyrighted-text continuation family showed partial replay behavior.
  • A prompt-injection benchmark using adversarial embedded content showed a partial success rate, meaning the model can sometimes follow malicious instructions embedded inside task content even when direct override attempts refuse.

Exact adversarial prompts and raw outputs are omitted from this Discord-safe copy.


Error-envelope / metadata summary

The endpoint’s error handling reveals implementation details such as exact cap messages, accepted reasoning-effort options, and pass-through instability on malformed roles. One raw error response contained a deploying-account identifier; that raw body is not included in this public copy.

Public conclusion: the error envelope is an OpenRouter-compatible proxy/wrapper artifact and should not be treated as direct upstream-model text.


What remains unproven

  • Exact checkpoint identity.
  • Official lab provenance or release plan.
  • Full system/developer prompt body.
  • Quantization status.
  • Exact attribution of every observed behavior to upstream weights vs provider wrapper vs gateway. The tokenizer/family finding is high confidence; some serving behaviors are clearly wrapper/gateway artifacts.

Methodology summary

  • Tokenizer differential against multiple candidate vocabularies.
  • Prompt-token accounting and cached-prefix subtraction.
  • Context ladder and needle retrieval checks.
  • Vision overhead checks.
  • Knowledge-bound probes with dated public facts.
  • Reasoning-effort and response-shape probes.
  • Redacted prompt-leak and injection probes.
  • Automated garak benchmark, summarized without raw adversarial prompts.

Public credits

  • OpenRouter public model catalog and OpenAI-compatible API behavior.
  • Hugging Face tokenizer files used for offline token counts only: MiniMax-Text-01, DeepSeek-V3, Qwen2.5-72B, GLM-4.5-Air, Llama-3-8B, and Kimi-K2.5.
  • NVIDIA garak v0.17.0 for automated benchmark coverage.
  • OWASP / MITRE prompt-leakage and instruction-hierarchy methodology references.
  • Recent 2026 prompt-leakage and instruction-hierarchy research used to guide the redacted probe categories.

MIT-style sharing is fine for this redacted write-up; raw measurements and probe scripts are intentionally excluded.

Sign up or log in to comment