Spaces:
Running
Running
Minette Kaunismäki commited on
Commit ·
d116f47
1
Parent(s): 1903aed
changing dataset name
Browse files
app.py
CHANGED
|
@@ -2665,7 +2665,7 @@ text_to_video_metric_ids = _metric_ids_for(
|
|
| 2665 |
datasets = [
|
| 2666 |
{
|
| 2667 |
"id": "text_to_video",
|
| 2668 |
-
"name": "
|
| 2669 |
"modality": "text_to_video",
|
| 2670 |
"data": text_to_video_df,
|
| 2671 |
"columns": text_to_video_display_columns,
|
|
|
|
| 2665 |
datasets = [
|
| 2666 |
{
|
| 2667 |
"id": "text_to_video",
|
| 2668 |
+
"name": "VBench-2.0 Dataset",
|
| 2669 |
"modality": "text_to_video",
|
| 2670 |
"data": text_to_video_df,
|
| 2671 |
"columns": text_to_video_display_columns,
|
ui.py
CHANGED
|
@@ -80,8 +80,8 @@ and price**. Each view is a **dataset** scored with a **metric**, written as
|
|
| 80 |
|
| 81 |
## How to read it
|
| 82 |
|
| 83 |
-
1. Pick a **type** (Text to Video, Video to Video, or Text to Image),
|
| 84 |
-
**dataset** and a **metric**.
|
| 85 |
2. **Leaderboards**: ranked by that metric. Price and generation time sit in
|
| 86 |
the same table when the source publishes them.
|
| 87 |
3. **Pareto plots**: mark models that are not beaten on both higher score
|
|
@@ -109,12 +109,12 @@ prompt suites, so samples are not shown.
|
|
| 109 |
|
| 110 |
## Current datasets
|
| 111 |
|
| 112 |
-
###
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
|
| 119 |
### Pruna Internal Video-Edit Benchmark
|
| 120 |
Pruna's internal video-to-video editing benchmark, collected by our
|
|
@@ -193,7 +193,7 @@ models *within* a Dataset | Metric view.
|
|
| 193 |
- **Prompt counts:** OneIG Alignment uses 100 anime, 100 human, and 99 object
|
| 194 |
prompts (299 total). Qwen Image Dataset uses 100 prompts sampled from the
|
| 195 |
1,000-prompt pool for roughly even coverage of its fine-grained (L3)
|
| 196 |
-
categories. The
|
| 197 |
generations per model. The Pruna Internal Video-Edit Benchmark uses 78
|
| 198 |
prompts across advertising, e-commerce, real estate, camera, lighting,
|
| 199 |
text, and related categories. Artificial Analysis and Arena AI use their
|
|
@@ -1385,11 +1385,11 @@ def _filter_row(datasets, metrics, default_dataset_id, default_metric_id=None):
|
|
| 1385 |
modality_dd = gr.Dropdown(
|
| 1386 |
choices=_modality_choices(datasets),
|
| 1387 |
value=default_modality,
|
| 1388 |
-
label="Type",
|
| 1389 |
type="value",
|
| 1390 |
filterable=False,
|
| 1391 |
scale=1,
|
| 1392 |
-
min_width=
|
| 1393 |
)
|
| 1394 |
dataset_dd = gr.Dropdown(
|
| 1395 |
choices=_dataset_choices(datasets, modality=default_modality),
|
|
@@ -1439,12 +1439,13 @@ def render_image_workspace(datasets, metrics, default_dataset_id, default_metric
|
|
| 1439 |
with gr.Column(elem_classes="workspace-filters") as filters_host:
|
| 1440 |
gr.Markdown(
|
| 1441 |
"<p class='filter-help'>"
|
| 1442 |
-
"Start with Type to switch
|
| 1443 |
-
"
|
| 1444 |
-
"
|
| 1445 |
-
"lists datasets and models we have generations
|
| 1446 |
-
"plots only lists datasets with price or
|
| 1447 |
-
"Search in Models, or leave it empty to
|
|
|
|
| 1448 |
"</p>",
|
| 1449 |
elem_classes="filter-help-host",
|
| 1450 |
)
|
|
|
|
| 80 |
|
| 81 |
## How to read it
|
| 82 |
|
| 83 |
+
1. Pick a **model type** (Text to Video, Video to Video, or Text to Image),
|
| 84 |
+
then a **dataset** and a **metric**.
|
| 85 |
2. **Leaderboards**: ranked by that metric. Price and generation time sit in
|
| 86 |
the same table when the source publishes them.
|
| 87 |
3. **Pareto plots**: mark models that are not beaten on both higher score
|
|
|
|
| 109 |
|
| 110 |
## Current datasets
|
| 111 |
|
| 112 |
+
### VBench-2.0 Dataset
|
| 113 |
+
VBench-2.0 prompts, comparing P-Video-2 variants with Fal-hosted models.
|
| 114 |
+
Quality is Datapoint Elo and Rapidata Elo from pairwise preference. Price
|
| 115 |
+
is USD per second of output video. Time per second of video is Fal wall
|
| 116 |
+
time, except Pruna models which use model execution time. Samples are not
|
| 117 |
+
shown for this dataset.
|
| 118 |
|
| 119 |
### Pruna Internal Video-Edit Benchmark
|
| 120 |
Pruna's internal video-to-video editing benchmark, collected by our
|
|
|
|
| 193 |
- **Prompt counts:** OneIG Alignment uses 100 anime, 100 human, and 99 object
|
| 194 |
prompts (299 total). Qwen Image Dataset uses 100 prompts sampled from the
|
| 195 |
1,000-prompt pool for roughly even coverage of its fine-grained (L3)
|
| 196 |
+
categories. The VBench-2.0 Dataset uses about 90
|
| 197 |
generations per model. The Pruna Internal Video-Edit Benchmark uses 78
|
| 198 |
prompts across advertising, e-commerce, real estate, camera, lighting,
|
| 199 |
text, and related categories. Artificial Analysis and Arena AI use their
|
|
|
|
| 1385 |
modality_dd = gr.Dropdown(
|
| 1386 |
choices=_modality_choices(datasets),
|
| 1387 |
value=default_modality,
|
| 1388 |
+
label="Model Type",
|
| 1389 |
type="value",
|
| 1390 |
filterable=False,
|
| 1391 |
scale=1,
|
| 1392 |
+
min_width=170,
|
| 1393 |
)
|
| 1394 |
dataset_dd = gr.Dropdown(
|
| 1395 |
choices=_dataset_choices(datasets, modality=default_modality),
|
|
|
|
| 1439 |
with gr.Column(elem_classes="workspace-filters") as filters_host:
|
| 1440 |
gr.Markdown(
|
| 1441 |
"<p class='filter-help'>"
|
| 1442 |
+
"Start with Model Type to switch modalities. The rest of "
|
| 1443 |
+
"the filters follow you across Leaderboards, Pareto plots, "
|
| 1444 |
+
"and Samples. "
|
| 1445 |
+
"Samples only lists datasets and models we have generations "
|
| 1446 |
+
"for; Pareto plots only lists datasets with price or "
|
| 1447 |
+
"generation time. Search in Models, or leave it empty to "
|
| 1448 |
+
"include every model."
|
| 1449 |
"</p>",
|
| 1450 |
elem_classes="filter-help-host",
|
| 1451 |
)
|