Instructions to use briaai/fibo-scene-analyzer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use briaai/fibo-scene-analyzer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="briaai/fibo-scene-analyzer") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("briaai/fibo-scene-analyzer") model = AutoModelForMultimodalLM.from_pretrained("briaai/fibo-scene-analyzer", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use briaai/fibo-scene-analyzer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "briaai/fibo-scene-analyzer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "briaai/fibo-scene-analyzer", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/briaai/fibo-scene-analyzer
- SGLang
How to use briaai/fibo-scene-analyzer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "briaai/fibo-scene-analyzer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "briaai/fibo-scene-analyzer", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "briaai/fibo-scene-analyzer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "briaai/fibo-scene-analyzer", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use briaai/fibo-scene-analyzer with Docker Model Runner:
docker model run hf.co/briaai/fibo-scene-analyzer
FIBO Scene Analyzer [dev]
FIBO Scene Analyzer [dev] was developed to cover a wide range of image understanding tasks using Bria's scene description format. A scene description holds a general description of the image (type, style, lighting, composition, color palette and more) and a tree of objects. Each object has a name, a description, its geometry (bounding boxes, polygon masks, body keypoints), its colors, its relationships to other objects and fields for its kind: people, animals, text and other objects.
The same model also does grounding (detection, masks, and body keypoints for a named target, or for all objects), OCR with line boxes, and text-to-description tasks: it expands a short prompt to a full scene description, and it refines or edits a scene description from an instruction.
We built FIBO Scene Analyzer from Qwen3.6-35B-A3B via multi-stage SFT on a custom synthetic dataset, then GRPO reinforcement learning with task-specific rewards, then a final short SFT stage. For synthetic data, we developed a 12-step workflow that combines computer vision tools, general and specialist models to extract all possible structured information from the images.
Supported tasks
| Task | Input | Prompt text | Answer |
|---|---|---|---|
caption |
image | <caption><DETAIL_LEVEL> |
SceneDescription |
detect |
image, optional target | <detect> or <detect>TARGET |
SceneObjectList: names, kinds, boxes |
mask |
image, optional target | <mask> or <mask>TARGET |
as detect, with polygon masks |
keypoints |
image, optional target | <keypoints> or <keypoints>TARGET |
people with boxes and named keypoints |
ocr |
image | <ocr> |
text objects: language, lines, line boxes |
describe |
image, object name | <describe>NAME |
one object with all its fields |
expand |
prompt, aspect ratio | <expand><DETAIL_LEVEL>\n<aspect_ratio>AR</aspect_ratio>\nPROMPT |
SceneDescription of a new image |
expand_object |
object name, optional scene description | <expand_object><DETAIL_LEVEL>NAME, or with \nSCENE_JSON after the name |
one object with all its fields |
refine |
instruction, scene description JSON | <refine><DETAIL_LEVEL>INSTRUCTION\nSCENE_JSON |
the changed SceneDescription |
edit |
1-5 images, instruction | <edit><DETAIL_LEVEL>\n<aspect_ratio>AR</aspect_ratio>\nINSTRUCTION |
SceneDescription of the edited image |
AR is the width divided by the height, rounded to two decimals.
With a scene description (scene_contexts in generate()), expand_object describes
the object so that it fits the scene, with relationships to the objects of the scene. The
scene must not contain the object itself.
SCENE_JSON is compacted JSON with coordinates and colors as
tokens. prompt_json() in scene_analyzer.py writes a SceneDescription in this form, also
after post_process().
For the grounding tasks, a named target that is not in the image gives an empty list.
Without a target, detect and mask return every object; clothing, footwear and
accessories are separate targets named "ITEM on PERSON".
Scene detail levels
| Detail level | Content |
|---|---|
general |
the general block only, without objects |
short |
a flat list of objects, with name, kind, description, location, position, relative size and text lines |
long |
as short, plus colors, details, relationships and the fields of each kind |
full |
the object tree with bounding boxes, relative distance, and body keypoints |
full_mask |
as full, with a polygon mask for most of the objects |
expand, refine and edit accept general, short, long and full.
expand_object accepts short, long, full and full_mask.
How to get started
Download the code files of this repository, then:
from PIL import Image
from scene_analyzer import SceneAnalyzerModel, SceneAnalyzerTask, DetailLevel
model = SceneAnalyzerModel("briaai/fibo-scene-analyzer") # Transformers, device_map="auto"
image = Image.open("photo.jpg").convert("RGB")
[description] = model.generate([image], SceneAnalyzerTask.CAPTION, detail_level=DetailLevel.FULL, num_samples=3)
description.post_process() # coordinate and color tokens to numbers
print(description.general.description)
for object in description.objects or []:
print(object.kind, object.name, object.bounding_boxes)
[people] = model.generate([image], SceneAnalyzerTask.DETECT, text_inputs=["person"])
[text] = model.generate([image], SceneAnalyzerTask.OCR)
To render an answer as an HTML page with the geometry drawn on the image:
from caption_visualizer import CaptionVisualizer
CaptionVisualizer().save("description.html", image, description)
Harnesses
To make it easier to work with the model and get the best results, we provide harness classes
based on the Pydantic scene description format.
The three harnesses have the same SceneAnalyzerModel.generate() interface and the same
task logic; only the engine differs.
| File | Engine | Use it for |
|---|---|---|
scene_analyzer.py |
Hugging Face Transformers | the smallest set of dependencies |
scene_analyzer_vllm.py |
vLLM, in the same process | batch processing; the output is constrained to the schema |
scene_analyzer_server.py |
an OpenAI-compatible server | a model served separately (vLLM, SGLang) |
What each harness does:
- It builds the prompt in the trained format. A prompt that the model was not trained
on, such as a detail level that a task does not accept, raises
ValueError. - It samples
num_samplesanswers and returns the best valid one. A scene description with the most objects wins; for grounding tasks, the most confident answer wins. An input without a valid answer is sampled again (retries), andNoneis returned if all attempts fail. - With
num_iterations, it runs revision rounds: the model sees its answer and adds the objects that it missed. This works for<caption>and the whole-image grounding tasks.
vLLM
from scene_analyzer_vllm import SceneAnalyzerModel
model = SceneAnalyzerModel("briaai/fibo-scene-analyzer", gpu_memory_utilization=0.9)
descriptions = model.generate(images, SceneAnalyzerTask.CAPTION, num_samples=3)
The model needs approximately 70 GB of GPU memory in bfloat16. On GPUs with less memory,
pass tensor_parallel_size.
OpenAI-compatible server
vllm serve briaai/fibo-scene-analyzer --max-model-len 32768 --limit-mm-per-prompt '{"image": 5}' \
--reasoning-parser deepseek_r1
from scene_analyzer_server import SceneAnalyzerModel
model = SceneAnalyzerModel("http://localhost:8000/v1")
descriptions = model.generate(images, SceneAnalyzerTask.CAPTION, num_samples=3, on_result=print)
The client sends standard chat-completion requests, and the chat template of the checkpoint renders the trained prompt. Two request fields outside the OpenAI schema are necessary for full function. vLLM and SGLang support both:
skip_special_tokens: false. The JSON structure and the coordinate and color values are special tokens, and servers remove special tokens from the output by default.chat_template_kwargs: {"thinking_effort": ...}selects a reasoning trace. Without it, the model answers directly.
Without a harness
The prompt is one user turn: the images, then the task tags. The assistant turn opens with an empty think block, and the model answers with JSON:
from transformers import AutoModelForImageTextToText, AutoProcessor
processor = AutoProcessor.from_pretrained("briaai/fibo-scene-analyzer")
model = AutoModelForImageTextToText.from_pretrained("briaai/fibo-scene-analyzer", dtype="auto", device_map="auto")
prompt = (
"<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|><caption><full><|im_end|>\n"
"<|im_start|>assistant\n<think>\n\n</think>\n\n"
)
inputs = processor(text=[prompt], images=[image], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=16384, do_sample=True, temperature=0.6, top_p=0.97, top_k=0)
answer = processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=False)
Decode with skip_special_tokens=False, then remove the final <|im_end|>.
Thinking
thinking_effort (minimal, low, medium or high) makes the model write a
reasoning trace before the answer. Direct answers are the default. Only caption,
expand, refine and edit are trained with traces, with these efforts:
| Detail level | Efforts |
|---|---|
general |
minimal |
short |
minimal, low |
long |
minimal, low, medium |
full, full_mask |
all |
Below is example of gradual prompt expansion using Chain-of-Thought.
Sampling
The harnesses sample with temperature 0.6 and top-p 0.97. Three samples per input
(num_samples=3) give the best scene descriptions; more samples do not help.
Scene description schema
data_model.py defines the schema with Pydantic. Pass
SceneDescription.model_json_schema() to a constrained decoder, or use
model_validate_json() to parse an answer. All fields except the ones marked
required are optional, and the model leaves out a field that does not apply.
Structure
SceneDescription the answer of `caption`, `expand`, `refine`, `edit`
├── general: GeneralDescription the image as a whole
└── objects: list[SceneObject] the objects in the image
├── Object.children objects that are part of the object or on it
├── Person.clothing / footwear / accessories: list[Object]
├── Person.eyes / hair / facial_hair: Object
├── Person.children, Animal.children: list[Object | Text]
└── Text.lines: list[TextLine]
SceneObjectList the answer of detect, mask, keypoints, ocr,
└── objects: list[SceneObject] describe and expand_object
SceneObject is one of Object, Person, Animal and Text, selected by kind.
Repeated objects are grouped: a row of chairs is one object with several boxes and
num_objects, not one object per chair.
General description
| Field | Type | Content |
|---|---|---|
user_prompt |
string | short scene description formatted as a prompt for text-to-image model |
description |
string | a short, complete description of the image |
image_type |
string | for example "photograph", "digital illustration", "graphic design", "3D render" |
style |
string | for example "photorealistic", "cinematic", "anime", "minimalist" |
medium |
string | for example "digital painting", "vector art" |
color_palette |
list of Color |
the main colors of the image |
aesthetic_score |
number, 0-10 | the estimated aesthetic quality |
atmosphere |
string | the mood of the image, for example "serene, majestic" |
setting |
string | the place, for example "city street at night", "studio backdrop" |
lighting |
string | for example "soft studio lighting", "bright daylight, direct sunlight" |
composition |
string | for example "centered, symmetrical", "asymmetrical balance" |
depth_of_field |
string | for example "shallow", "deep" |
viewpoint |
string | for example "eye-level", "front view" |
intent |
string | the purpose of the image, for example "promotional movie poster" |
artifacts |
list of strings | visual defects of the whole image |
camera_model, lens_model |
string | the camera and the lens, when they can be estimated |
focal_length, aperture, f_number, shutter_speed, exposure_time_ms, iso |
number | the exposure settings, when they can be estimated |
Fields of every object
| Field | Type | Content | Detail levels |
|---|---|---|---|
kind |
object, person, animal, text |
the class of the object | all |
name |
string | a short name, unique in the scene description, for example "man in a tuxedo" | all |
description |
string | a longer description of the object | all |
num_objects |
integer | the number of instances, when the object is a group | all |
location |
string | the place in the frame: for example "top-left corner", "left of center", "bottom center", "full frame" | short, long |
position |
foreground, midground, background |
the depth band | short, long |
relative_size |
small, medium, large |
the area of the object against the image | short, long |
colors |
list of Color |
the main colors of the object | long, full, full_mask |
details |
list of strings | more visual details | long, full, full_mask |
relationships |
list of Relationship |
relations to other objects | long, full, full_mask |
blur |
string | for example "slightly blurred", "out of focus", "motion blur" | long, full, full_mask |
artifacts |
list of strings | visual defects of the object | long, full, full_mask |
bounding_boxes |
list of BoundingBox |
one box per instance | full, full_mask |
relative_distance |
number, 0-100 | the distance from the camera: 0 is at the camera, 100 at the far end of the scene | full, full_mask |
mask |
ObjectMask |
the outline of the object | full_mask |
location, position and relative_size replace the geometry in the lower detail levels.
data_model.py calculates them from bounding_boxes and relative_distance
(make_long(), make_short()), so a full scene description converts to the lower detail
levels without the model.
Fields by kind
object
| Field | Type | Content |
|---|---|---|
material |
string | for example "metal", "wood", "organic plant matter" |
texture |
string | for example "smooth", "rough", "matte" |
transparency |
number, 0-100 | 0 is fully transparent, 100 is fully opaque |
children |
list of objects | objects that are part of the object or on it |
person
| Field | Type | Content |
|---|---|---|
age, gender, ethnicity |
string | visual estimates, for example "20s", "female"; "unknown" when they cannot be seen |
skin_color |
Color |
the skin tone, on the Monk skin tone scale |
activity |
string | what the person does |
expression |
string | for example "neutral", "focused", "not visible" |
eyes, hair, facial_hair |
Object |
the face parts, with their own colors and details |
mouth, nose, ears, body, left_hand, right_hand, left_leg, right_leg |
string | descriptions of the body parts |
face_bbox |
BoundingBox |
the box of the face |
keypoints |
list of BodyKeypoint |
the visible body, face, hand and foot points |
clothing, footwear, accessories |
list of Object |
what the person wears or carries |
children |
list of Object or Text |
other objects on the person, such as a print on a shirt |
animal
| Field | Type | Content |
|---|---|---|
head, body, legs, hands, tail, wings |
string | descriptions of the body parts |
children |
list of Object or Text |
objects on the animal, such as a collar |
text
| Field | Type | Content |
|---|---|---|
lines |
list of TextLine, required |
the lines of the text, each with text and its box |
language |
string | the language code, for example "en", "ja", "de" |
font_family |
string | for example "sans-serif", "serif", "monospace", "cursive" |
font_description |
string | a description of the typeface |
font_weight |
"100" to "900" | the font weight |
font_style |
normal, italic, oblique |
|
text_case |
lowercase, uppercase, title, sentence, mixed |
|
text_orientation |
horizontal, vertical |
|
text_alignment |
left, center, right, justify |
|
text_decoration |
none, underline, line-through, overline, shadow |
|
rotation |
number, -180 to 180 | the rotation of the text, in degrees |
A TextLine has the fields of every object and text (required). Its kind is
base_object. In full scene descriptions, each line has its own bounding_boxes.
Field presence by detail level
| Detail level | Objects | Geometry | Kind fields |
|---|---|---|---|
general |
none | none | none |
short |
a flat list | location, position, relative_size |
age, gender, ethnicity; the text lines |
long |
a flat list | location, position, relative_size |
all, without face_bbox and keypoints |
full |
the tree | boxes, relative_distance, keypoints, face box, line boxes |
all |
full_mask |
the tree | as full, and mask |
all |
In short and long, the children of each object move to the top-level list. The
wearables stay under their person in long, and short leaves them out.
Geometry and colors
| Type | Fields | Content |
|---|---|---|
Point |
x, y, required |
percentages of the image width and height, 0-100 with one decimal, from the top-left corner |
BoundingBox |
top_left, bottom_right: Point, required |
an axis-aligned box |
ObjectMask |
polygons: list of lists of numbers |
one or more polygons, each a flat list [x1, y1, x2, y2, ...] in Point units |
BodyKeypoint |
name, point, required |
a named point (see the names below) |
Color |
red, green, blue: 0-255, required; name |
a color, with a color name such as "Madder Lake" |
Relationship |
to, relationship, required |
the name of the other object, and the relation, for example "holding" |
ObjectMask.decode(width, height) rasterizes the polygons to a boolean array.
In the raw answer, coordinates and color channels are tokens: "<coord_12.3>" and
"<color_200>". post_process() converts them to numbers in place.
The keypoint names ("left" and "right" are the person's own sides):
- Body: nose, neck, and left/right eye, ear, shoulder, elbow, wrist, hip, knee and ankle.
- Face: chin, nose tip, upper lip center, lower lip center, and left/right iris, eyebrow inner, eyebrow outer and mouth corner.
- Hands: left/right thumb tip, index fingertip, middle fingertip, ring fingertip and pinky fingertip.
- Feet: left/right heel, big toe and small toe.
Examples
The images show answers of FIBO Scene Analyzer [dev]. The images for the general caption, OCR, detection, masks, keypoints, prompt expansion and object expansion use the harness defaults: direct answers, temperature 0.6, top-p 0.97 and three samples per input. Their input photos are from Wikimedia Commons and Flickr, and the photo credits are at the end of this section.
Full caption
<caption><full> gives the complete scene description: the general description and every
object with its geometry, colors and fields. Each image shows the input, the description
drawn on the image and a part of the JSON.
More examples (5)
General caption
<caption><general> gives only the general description: image type, style, medium, color palette, lighting, composition and a short description.
More examples (4)
OCR
<ocr> reads each line of text, gives a box for each line, and tags each text object with its language. The model was primarily
trained for English, but it is also capable of understanding other languages.
More examples (6)
Named object detection
Each number is one <detect>NAME prompt. The last name in each list is not in the image, and its answer is an empty list.
More examples (4)
Masks
<mask> without a target gives a polygon mask for each object. Clothing, footwear and accessories are separate objects, named "ITEM on PERSON".
More examples (4)
Keypoints
<keypoints> without a target gives each person with a box and named keypoints on the body, feet, hands and face.
More examples (4)
Prompt expansion
<expand><full> writes the scene description of a new image from a short prompt and an aspect ratio. The images show its layout: the box of each object, the boxes of clothing and footwear, and the keypoints of each person.
More examples (4)
Object expansion
<expand_object><full> describes one new object for a scene. With the scene description of an
<expand> answer as context, it places the object in the scene and gives its relationships
to the objects of the scene. The left side of each image shows the prompt and its scene. The
right side shows the object prompt and the scene with the new object (purple) and the scene
objects that it relates to (teal).
More examples (5)
Citations
@article{kachlon2026bbq,
title={BBQ-to-Image: Numeric Bounding Box and Qolor Control in Large-Scale Text-to-Image Models},
author={Eliran Kachlon and Alexander Visheratin and Nimrod Sarid and Tal Hacham and Eyal Gutflaish and Saar Huberman and Hezi Zisman and David Ruppin and Ron Mokady},
journal={arXiv preprint arXiv:2602.20672},
year={2026}
}
- Downloads last month
- 98
Model tree for briaai/fibo-scene-analyzer
Base model
Qwen/Qwen3.6-35B-A3B