FIBO Scene Analyzer [dev]

Demo on Hugging Face Spaces Interactive examples gallery

FIBO Scene Analyzer [dev] was developed to cover a wide range of image understanding tasks using Bria's scene description format. A scene description holds a general description of the image (type, style, lighting, composition, color palette and more) and a tree of objects. Each object has a name, a description, its geometry (bounding boxes, polygon masks, body keypoints), its colors, its relationships to other objects and fields for its kind: people, animals, text and other objects.

The same model also does grounding (detection, masks, and body keypoints for a named target, or for all objects), OCR with line boxes, and text-to-description tasks: it expands a short prompt to a full scene description, and it refines or edits a scene description from an instruction.

We built FIBO Scene Analyzer from Qwen3.6-35B-A3B via multi-stage SFT on a custom synthetic dataset, then GRPO reinforcement learning with task-specific rewards, then a final short SFT stage. For synthetic data, we developed a 12-step workflow that combines computer vision tools, general and specialist models to extract all possible structured information from the images.

Supported tasks

Task Input Prompt text Answer
caption image <caption><DETAIL_LEVEL> SceneDescription
detect image, optional target <detect> or <detect>TARGET SceneObjectList: names, kinds, boxes
mask image, optional target <mask> or <mask>TARGET as detect, with polygon masks
keypoints image, optional target <keypoints> or <keypoints>TARGET people with boxes and named keypoints
ocr image <ocr> text objects: language, lines, line boxes
describe image, object name <describe>NAME one object with all its fields
expand prompt, aspect ratio <expand><DETAIL_LEVEL>\n<aspect_ratio>AR</aspect_ratio>\nPROMPT SceneDescription of a new image
expand_object object name, optional scene description <expand_object><DETAIL_LEVEL>NAME, or with \nSCENE_JSON after the name one object with all its fields
refine instruction, scene description JSON <refine><DETAIL_LEVEL>INSTRUCTION\nSCENE_JSON the changed SceneDescription
edit 1-5 images, instruction <edit><DETAIL_LEVEL>\n<aspect_ratio>AR</aspect_ratio>\nINSTRUCTION SceneDescription of the edited image

AR is the width divided by the height, rounded to two decimals.

With a scene description (scene_contexts in generate()), expand_object describes the object so that it fits the scene, with relationships to the objects of the scene. The scene must not contain the object itself.

SCENE_JSON is compacted JSON with coordinates and colors as tokens. prompt_json() in scene_analyzer.py writes a SceneDescription in this form, also after post_process().

For the grounding tasks, a named target that is not in the image gives an empty list. Without a target, detect and mask return every object; clothing, footwear and accessories are separate targets named "ITEM on PERSON".

Scene detail levels

Detail level Content
general the general block only, without objects
short a flat list of objects, with name, kind, description, location, position, relative size and text lines
long as short, plus colors, details, relationships and the fields of each kind
full the object tree with bounding boxes, relative distance, and body keypoints
full_mask as full, with a polygon mask for most of the objects

expand, refine and edit accept general, short, long and full. expand_object accepts short, long, full and full_mask.

How to get started

Download the code files of this repository, then:

from PIL import Image
from scene_analyzer import SceneAnalyzerModel, SceneAnalyzerTask, DetailLevel

model = SceneAnalyzerModel("briaai/fibo-scene-analyzer")  # Transformers, device_map="auto"
image = Image.open("photo.jpg").convert("RGB")

[description] = model.generate([image], SceneAnalyzerTask.CAPTION, detail_level=DetailLevel.FULL, num_samples=3)
description.post_process()  # coordinate and color tokens to numbers
print(description.general.description)
for object in description.objects or []:
    print(object.kind, object.name, object.bounding_boxes)

[people] = model.generate([image], SceneAnalyzerTask.DETECT, text_inputs=["person"])
[text] = model.generate([image], SceneAnalyzerTask.OCR)

To render an answer as an HTML page with the geometry drawn on the image:

from caption_visualizer import CaptionVisualizer

CaptionVisualizer().save("description.html", image, description)

Harnesses

To make it easier to work with the model and get the best results, we provide harness classes based on the Pydantic scene description format. The three harnesses have the same SceneAnalyzerModel.generate() interface and the same task logic; only the engine differs.

File Engine Use it for
scene_analyzer.py Hugging Face Transformers the smallest set of dependencies
scene_analyzer_vllm.py vLLM, in the same process batch processing; the output is constrained to the schema
scene_analyzer_server.py an OpenAI-compatible server a model served separately (vLLM, SGLang)

What each harness does:

  • It builds the prompt in the trained format. A prompt that the model was not trained on, such as a detail level that a task does not accept, raises ValueError.
  • It samples num_samples answers and returns the best valid one. A scene description with the most objects wins; for grounding tasks, the most confident answer wins. An input without a valid answer is sampled again (retries), and None is returned if all attempts fail.
  • With num_iterations, it runs revision rounds: the model sees its answer and adds the objects that it missed. This works for <caption> and the whole-image grounding tasks.

vLLM

from scene_analyzer_vllm import SceneAnalyzerModel

model = SceneAnalyzerModel("briaai/fibo-scene-analyzer", gpu_memory_utilization=0.9)
descriptions = model.generate(images, SceneAnalyzerTask.CAPTION, num_samples=3)

The model needs approximately 70 GB of GPU memory in bfloat16. On GPUs with less memory, pass tensor_parallel_size.

OpenAI-compatible server

vllm serve briaai/fibo-scene-analyzer --max-model-len 32768 --limit-mm-per-prompt '{"image": 5}' \
    --reasoning-parser deepseek_r1
from scene_analyzer_server import SceneAnalyzerModel

model = SceneAnalyzerModel("http://localhost:8000/v1")
descriptions = model.generate(images, SceneAnalyzerTask.CAPTION, num_samples=3, on_result=print)

The client sends standard chat-completion requests, and the chat template of the checkpoint renders the trained prompt. Two request fields outside the OpenAI schema are necessary for full function. vLLM and SGLang support both:

  • skip_special_tokens: false. The JSON structure and the coordinate and color values are special tokens, and servers remove special tokens from the output by default.
  • chat_template_kwargs: {"thinking_effort": ...} selects a reasoning trace. Without it, the model answers directly.

Without a harness

The prompt is one user turn: the images, then the task tags. The assistant turn opens with an empty think block, and the model answers with JSON:

from transformers import AutoModelForImageTextToText, AutoProcessor

processor = AutoProcessor.from_pretrained("briaai/fibo-scene-analyzer")
model = AutoModelForImageTextToText.from_pretrained("briaai/fibo-scene-analyzer", dtype="auto", device_map="auto")
prompt = (
    "<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|><caption><full><|im_end|>\n"
    "<|im_start|>assistant\n<think>\n\n</think>\n\n"
)
inputs = processor(text=[prompt], images=[image], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=16384, do_sample=True, temperature=0.6, top_p=0.97, top_k=0)
answer = processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=False)

Decode with skip_special_tokens=False, then remove the final <|im_end|>.

Thinking

thinking_effort (minimal, low, medium or high) makes the model write a reasoning trace before the answer. Direct answers are the default. Only caption, expand, refine and edit are trained with traces, with these efforts:

Detail level Efforts
general minimal
short minimal, low
long minimal, low, medium
full, full_mask all

Below is example of gradual prompt expansion using Chain-of-Thought.

Sampling

The harnesses sample with temperature 0.6 and top-p 0.97. Three samples per input (num_samples=3) give the best scene descriptions; more samples do not help.

Scene description schema

data_model.py defines the schema with Pydantic. Pass SceneDescription.model_json_schema() to a constrained decoder, or use model_validate_json() to parse an answer. All fields except the ones marked required are optional, and the model leaves out a field that does not apply.

Structure

SceneDescription                       the answer of `caption`, `expand`, `refine`, `edit`
├── general: GeneralDescription        the image as a whole
└── objects: list[SceneObject]         the objects in the image
    ├── Object.children                objects that are part of the object or on it
    ├── Person.clothing / footwear / accessories: list[Object]
    ├── Person.eyes / hair / facial_hair: Object
    ├── Person.children, Animal.children: list[Object | Text]
    └── Text.lines: list[TextLine]

SceneObjectList                        the answer of detect, mask, keypoints, ocr,
└── objects: list[SceneObject]         describe and expand_object

SceneObject is one of Object, Person, Animal and Text, selected by kind. Repeated objects are grouped: a row of chairs is one object with several boxes and num_objects, not one object per chair.

General description

Field Type Content
user_prompt string short scene description formatted as a prompt for text-to-image model
description string a short, complete description of the image
image_type string for example "photograph", "digital illustration", "graphic design", "3D render"
style string for example "photorealistic", "cinematic", "anime", "minimalist"
medium string for example "digital painting", "vector art"
color_palette list of Color the main colors of the image
aesthetic_score number, 0-10 the estimated aesthetic quality
atmosphere string the mood of the image, for example "serene, majestic"
setting string the place, for example "city street at night", "studio backdrop"
lighting string for example "soft studio lighting", "bright daylight, direct sunlight"
composition string for example "centered, symmetrical", "asymmetrical balance"
depth_of_field string for example "shallow", "deep"
viewpoint string for example "eye-level", "front view"
intent string the purpose of the image, for example "promotional movie poster"
artifacts list of strings visual defects of the whole image
camera_model, lens_model string the camera and the lens, when they can be estimated
focal_length, aperture, f_number, shutter_speed, exposure_time_ms, iso number the exposure settings, when they can be estimated

Fields of every object

Field Type Content Detail levels
kind object, person, animal, text the class of the object all
name string a short name, unique in the scene description, for example "man in a tuxedo" all
description string a longer description of the object all
num_objects integer the number of instances, when the object is a group all
location string the place in the frame: for example "top-left corner", "left of center", "bottom center", "full frame" short, long
position foreground, midground, background the depth band short, long
relative_size small, medium, large the area of the object against the image short, long
colors list of Color the main colors of the object long, full, full_mask
details list of strings more visual details long, full, full_mask
relationships list of Relationship relations to other objects long, full, full_mask
blur string for example "slightly blurred", "out of focus", "motion blur" long, full, full_mask
artifacts list of strings visual defects of the object long, full, full_mask
bounding_boxes list of BoundingBox one box per instance full, full_mask
relative_distance number, 0-100 the distance from the camera: 0 is at the camera, 100 at the far end of the scene full, full_mask
mask ObjectMask the outline of the object full_mask

location, position and relative_size replace the geometry in the lower detail levels. data_model.py calculates them from bounding_boxes and relative_distance (make_long(), make_short()), so a full scene description converts to the lower detail levels without the model.

Fields by kind

object

Field Type Content
material string for example "metal", "wood", "organic plant matter"
texture string for example "smooth", "rough", "matte"
transparency number, 0-100 0 is fully transparent, 100 is fully opaque
children list of objects objects that are part of the object or on it

person

Field Type Content
age, gender, ethnicity string visual estimates, for example "20s", "female"; "unknown" when they cannot be seen
skin_color Color the skin tone, on the Monk skin tone scale
activity string what the person does
expression string for example "neutral", "focused", "not visible"
eyes, hair, facial_hair Object the face parts, with their own colors and details
mouth, nose, ears, body, left_hand, right_hand, left_leg, right_leg string descriptions of the body parts
face_bbox BoundingBox the box of the face
keypoints list of BodyKeypoint the visible body, face, hand and foot points
clothing, footwear, accessories list of Object what the person wears or carries
children list of Object or Text other objects on the person, such as a print on a shirt

animal

Field Type Content
head, body, legs, hands, tail, wings string descriptions of the body parts
children list of Object or Text objects on the animal, such as a collar

text

Field Type Content
lines list of TextLine, required the lines of the text, each with text and its box
language string the language code, for example "en", "ja", "de"
font_family string for example "sans-serif", "serif", "monospace", "cursive"
font_description string a description of the typeface
font_weight "100" to "900" the font weight
font_style normal, italic, oblique
text_case lowercase, uppercase, title, sentence, mixed
text_orientation horizontal, vertical
text_alignment left, center, right, justify
text_decoration none, underline, line-through, overline, shadow
rotation number, -180 to 180 the rotation of the text, in degrees

A TextLine has the fields of every object and text (required). Its kind is base_object. In full scene descriptions, each line has its own bounding_boxes.

Field presence by detail level

Detail level Objects Geometry Kind fields
general none none none
short a flat list location, position, relative_size age, gender, ethnicity; the text lines
long a flat list location, position, relative_size all, without face_bbox and keypoints
full the tree boxes, relative_distance, keypoints, face box, line boxes all
full_mask the tree as full, and mask all

In short and long, the children of each object move to the top-level list. The wearables stay under their person in long, and short leaves them out.

Geometry and colors

Type Fields Content
Point x, y, required percentages of the image width and height, 0-100 with one decimal, from the top-left corner
BoundingBox top_left, bottom_right: Point, required an axis-aligned box
ObjectMask polygons: list of lists of numbers one or more polygons, each a flat list [x1, y1, x2, y2, ...] in Point units
BodyKeypoint name, point, required a named point (see the names below)
Color red, green, blue: 0-255, required; name a color, with a color name such as "Madder Lake"
Relationship to, relationship, required the name of the other object, and the relation, for example "holding"

ObjectMask.decode(width, height) rasterizes the polygons to a boolean array.

In the raw answer, coordinates and color channels are tokens: "<coord_12.3>" and "<color_200>". post_process() converts them to numbers in place.

The keypoint names ("left" and "right" are the person's own sides):

  • Body: nose, neck, and left/right eye, ear, shoulder, elbow, wrist, hip, knee and ankle.
  • Face: chin, nose tip, upper lip center, lower lip center, and left/right iris, eyebrow inner, eyebrow outer and mouth corner.
  • Hands: left/right thumb tip, index fingertip, middle fingertip, ring fingertip and pinky fingertip.
  • Feet: left/right heel, big toe and small toe.

Examples

The images show answers of FIBO Scene Analyzer [dev]. The images for the general caption, OCR, detection, masks, keypoints, prompt expansion and object expansion use the harness defaults: direct answers, temperature 0.6, top-p 0.97 and three samples per input. Their input photos are from Wikimedia Commons and Flickr, and the photo credits are at the end of this section.

Full caption

<caption><full> gives the complete scene description: the general description and every object with its geometry, colors and fields. Each image shows the input, the description drawn on the image and a part of the JSON.

FIBO Scene Analyzer: an image, its scene description drawn on the image, and the JSON

More examples (5)

FIBO Scene Analyzer on a poster: objects, a person with keypoints, and the text lines read in place

FIBO Scene Analyzer on a dance photo: the dancer's box, face box and full-body keypoints

FIBO Scene Analyzer on a workshop photo: two people with keypoints, a workbench and shelves

FIBO Scene Analyzer on a busy illustration: ten people and a landmark, each with its box

FIBO Scene Analyzer on a Bauhaus poster: a stylized figure, color blocks and the text lines

General caption

<caption><general> gives only the general description: image type, style, medium, color palette, lighting, composition and a short description.

General caption of an Impressionist painting of a garden with sunflowers

More examples (4)

General caption of a studio photo of two lemons and a lime

General caption of a misty bog at sunrise

General caption of a street vendor on a sidewalk

General caption of a flat book illustration of an alley

OCR

<ocr> reads each line of text, gives a box for each line, and tags each text object with its language. The model was primarily trained for English, but it is also capable of understanding other languages.

OCR on six signs: Japanese and English, Chinese, Russian, Greek, Vietnamese and Spanish, with line boxes and language tags

More examples (6)

OCR on a Japanese station sign: the station names in kanji and in romaji

OCR on a Chinese fire extinguisher box: Chinese and English lines

OCR on a Russian memorial plaque: seven lines in capitals

OCR on an Italian marble plaque, with the old spellings QVI FV and COMVNE

OCR on an English blue plaque, including the curved lines at the top and the bottom

OCR on an English wheel-clamp warning sign

Named object detection

Each number is one <detect>NAME prompt. The last name in each list is not in the image, and its answer is an empty list.

Named detection on a pepper market: 21 prompts, one box each, including all 16 bowls

More examples (4)

Named detection in an old office: 23 prompts, from the windows to the stamps on the desk

Named detection in a hotel study: 25 prompts, including the two ceiling spotlights

Named detection of six nesting dolls by size and position

Named detection in a spa bathroom: 15 prompts

Masks

<mask> without a target gives a polygon mask for each object. Clothing, footwear and accessories are separate objects, named "ITEM on PERSON".

Masks of a cat, a leather couch and a rug

More examples (4)

Masks of a woman, her headscarf, her jacket and the hillside behind her

Masks of a scooter, its top case and its shadow, a wall, a door and the ground

Masks of two cupcakes, their plates, a cup, a saucer, a spoon and the table

Masks of a glass vase, a glass pumpkin, a glass pear, the tabletop and the wall

Keypoints

<keypoints> without a target gives each person with a box and named keypoints on the body, feet, hands and face.

Keypoints of a Bharatanatyam dancer with one leg raised

More examples (4)

Keypoints of two acrobats on a vertical pole, one upside down

Keypoints of a motocross rider in the air

Keypoints of a figure-skating pair in a death spiral

Keypoints of a guitarist, including the fingertips on the strings

Prompt expansion

<expand><full> writes the scene description of a new image from a short prompt and an aspect ratio. The images show its layout: the box of each object, the boxes of clothing and footwear, and the keypoints of each person.

Layout from the prompt “A family of four walking hand in hand along a beach.” at 16:9

More examples (4)

Layout from a prompt for a high karate kick in a dojo, with the instructor watching, at 3:2

Layout from a prompt for a ballerina in an arabesque on an empty stage, at 4:5

Layout from a prompt for two people riding bicycles side by side on a country road, at 16:9

Layout from a prompt for a boy holding a kitten in a garden, at 4:5

Object expansion

<expand_object><full> describes one new object for a scene. With the scene description of an <expand> answer as context, it places the object in the scene and gives its relationships to the objects of the scene. The left side of each image shows the prompt and its scene. The right side shows the object prompt and the scene with the new object (purple) and the scene objects that it relates to (teal).

Object expansion in a music room: a girl sitting on the bench and playing the piano, with her keypoints

More examples (5)

Object expansion in a garden: an old man reading a newspaper on the wooden bench

Object expansion in a meadow: a girl sitting on the wooden swing

Object expansion in a dining room: a vase of flowers in the middle of the table

Object expansion in a reading nook: a reading lamp next to the armchair

Object expansion in the same reading nook: a cat sleeping in the armchair

Citations

@article{kachlon2026bbq,
  title={BBQ-to-Image: Numeric Bounding Box and Qolor Control in Large-Scale Text-to-Image Models},
  author={Eliran Kachlon and Alexander Visheratin and Nimrod Sarid and Tal Hacham and Eyal Gutflaish and Saar Huberman and Hezi Zisman and David Ruppin and Ron Mokady},
  journal={arXiv preprint arXiv:2602.20672},
  year={2026}
}
Downloads last month
98
Safetensors
Model size
36B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for briaai/fibo-scene-analyzer

Finetuned
(341)
this model

Space using briaai/fibo-scene-analyzer 1

Paper for briaai/fibo-scene-analyzer