ec2eat Laya endpoint wrapper

This repository is a custom Hugging Face Inference Toolkit handler, not a standard Transformers classification checkpoint. It downloads only convaiinnovations/laya-multilingual revision e4e9ddf21a7b1903b7acffd8814ad4307bf63a67 at startup and runs Laya 0.3.20 on CPU. Upstream weights are Apache-2.0; they are not redistributed in this wrapper.

Current status

Candidate serving recipe. Five local validation tests passed. Real CPU inference passed English and Traditional Chinese two-candidate examples; both ranked the matching lunch first. These examples do not establish general ranking quality. Local unit tests are separate from real local model tests and from HF deployment validation. The application still exports a null installed recipe and falls back deterministically. Uploading these files does not enable it.

Deploy in steps

  1. Create your own model repository, proposed name ylm182/ec2eat-laya-serving. Upload handler.py, requirements.txt, and this README.md at its root. The browser upload uses your signed-in account; do not widen the production inference token to write access. Do not upload tokens, caches or virtual environments.
  2. Record the new repository's full commit hash. Use your wrapper repository for the endpoint, not the upstream weights repository. The endpoint commit pins the handler; the handler separately pins the upstream weights.
  3. Select an Inference Toolkit engine supporting handler.py (the legacy Default engine detects it). Check initialization logs to confirm custom handler discovery. If Default maps to another engine, stop and select Inference Toolkit or use its documented container; do not substitute vLLM/TGI or a stock classification pipeline.
  4. Starting hardware: AWS Intel Sapphire Rapids CPU, 4 vCPUs / 8 GB. Region depends on availability; the user's draft offers N. Virginia. Measure latency from Taiwan. Use owner-token-only authentication (labelled Private in the current UI), not public access. Leave AWS PrivateLink off. Replicas 0โ€“1, idle timeout 60 minutes. Confirm the displayed hourly cost before creation. This size is an estimate until Linux startup memory and latency are measured.
  5. Leave server command/arguments empty. No Gemini, Calendar, Places or Firebase credentials belong here. The checkpoint is public. First startup downloads weights.
  6. After deployment, record actual image digest/runtime and send a synthetic request from smoke.py to the endpoint root /, with Bearer HF authentication. Verify an anonymous request fails. Save sanitized requests/responses plus metadata.
  7. Only then implement the application's VerifiedLayaRecipe, with fixture digest, handler/runtime version and checkpoint revision. Run the existing live contract checklist in tests/fixtures/laya/README.md before enabling integration.

Contract

POST the JSON envelope {"inputs":{"state":"...","candidates":[{"id":"a","description":"..."}]}}. State is minimized preference/context text, never raw Calendar data. Production encoding must explain dimension polarity, preserve explicit/neutral/unknown states, keep priors weak, and respect the existing Places model-input flag.

1โ€“10 unique candidate IDs, 16,000-byte request maximum, at most 700 state tokens including candidate descriptions. Over-limit requests fail instead of truncating. One fixed choice question returns scores for every candidate. Letter aliases keep IDs out of the question head; the response restores original IDs. Raw probabilities are normalized only to correct upstream four-decimal rounding. They are ranking weights, not calibrated enjoyment probabilities; confidence is null.

The app's 2-second timeout remains unchanged. Slow/cold model responses must fall back. Local Apple CPU timing does not establish AWS latency. Scale-up, warm-up, scheduler authorization, idle scaling and Linux dependency compatibility require real deployment tests.

Local verification

Use Python 3.12 in an isolated virtual environment, install requirements.txt, then:

python -m unittest discover -s tests -v
python smoke.py

The smoke test downloads real public weights and creates local-smoke.json using fictional data only. local-validation-requirements.txt records the locally resolved packages; it is not a claim that HF's base image has been tested.

Sources checked 2026-09-25:

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support