Image-Text-to-Text
LiteRT-LM
LiteRT
English
decision-model
vision
structured-output
hybrid
gated-deltanet
Instructions to use litert-community/decider-2b-vision-LiteRT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/decider-2b-vision-LiteRT with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli # A single .litertlm file in the repo is picked automatically; otherwise the CLI asks which one to run # (or pass its name right after the repo id). litert-lm run \ --from-huggingface-repo=litert-community/decider-2b-vision-LiteRT \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/decider-2b-vision-LiteRT with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
File size: 3,638 Bytes
6c024e9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 | """The LiteRT-LM runtime path: one image and one question in, the answer letter out.
pip install litert-lm==0.17.1 tokenizers pillow
python -B reference/runtime_example.py --bundle decider-2b-vision_fp16-int8vocab.litertlm --image frame.png \
--context "You play Pong (Atari) and control the right paddle. ..." \
--question "What should you do right now?" --options "move paddle up" "move paddle down" "stay"
Scope of this path (what was checked): the image comes first, there is exactly ONE question, and the answer is the
first generated token with greedy decoding. The runtime returns text, not option probabilities, and its greedy token
is the most likely token over the whole vocabulary, not only over the option letters (the two agree on every
single-question published fixture row, but that is not guaranteed). For probabilities, several questions in one
request, or a text-only request, use decider_litert.py.
The text is built with the upstream prompt code (vendored), and the image is resized to 256 x 256 with PIL BICUBIC
before it is handed to the runtime, as in the conversion checks.
"""
import argparse
import os
import sys
import tempfile
import time
from pathlib import Path
HERE = Path(__file__).resolve().parent
sys.path.insert(0, str(HERE))
from bundle_cache import read_tokenizer_json # noqa: E402
from decider_litert import HFTokenizer, request_text, image_256 # noqa: E402
LETTERS = 'ABCDEFGHIJ'
def main():
ap = argparse.ArgumentParser()
ap.add_argument('--bundle', required=True)
ap.add_argument('--image', required=True)
ap.add_argument('--context', required=True)
ap.add_argument('--question', required=True)
ap.add_argument('--options', nargs='+', required=True)
ap.add_argument('--backend', choices=('cpu', 'gpu'), default='cpu')
ap.add_argument('--cache-dir', default=None, help='runtime cache folder (default: next to the bundle); ":nocache" for none')
args = ap.parse_args()
import litert_lm
from litert_lm import Backend, Engine, SamplerConfig
tok = HFTokenizer(json_text=read_tokenizer_json(args.bundle))
text, nopts, _ = request_text(tok, args.context, [dict(question=args.question, options=args.options)])
with tempfile.TemporaryDirectory() as tmp:
png = os.path.join(tmp, 'image_256.png')
image_256(args.image).save(png, format='PNG')
backend = Backend.CPU() if args.backend == 'cpu' else Backend.GPU()
t0 = time.monotonic()
engine = Engine(args.bundle, backend=backend, vision_backend=backend, max_num_tokens=4096, max_num_images=1,
cache_dir=args.cache_dir)
t1 = time.monotonic()
conv = engine.create_conversation(sampler_config=SamplerConfig(top_k=1), max_output_tokens=1)
message = dict(role='user', content=[dict(type='image', path=png), dict(type='text', text=text)])
first = ''
for chunk in conv.send_message_async(message):
c = chunk.get('content', '')
if isinstance(c, list):
c = ''.join(p.get('text', '') for p in c if isinstance(p, dict))
first += c
t2 = time.monotonic()
conv.close()
engine.close()
letter = first.strip()
if letter in LETTERS[:nopts[0]]:
print(f'answer: ({letter}) {args.options[LETTERS.index(letter)]}')
else:
print(f'the greedy token {first!r} is not one of the option letters; read probabilities with decider_litert.py')
print(f'engine {t1 - t0:.1f} s, request {t2 - t1:.2f} s', file=sys.stderr)
if __name__ == '__main__':
main()
|