Minette Kaunismäki commited on
Commit
d116f47
·
1 Parent(s): 1903aed

changing dataset name

Browse files
Files changed (2) hide show
  1. app.py +1 -1
  2. ui.py +18 -17
app.py CHANGED
@@ -2665,7 +2665,7 @@ text_to_video_metric_ids = _metric_ids_for(
2665
  datasets = [
2666
  {
2667
  "id": "text_to_video",
2668
- "name": "Pruna Internal Text-to-Video Benchmark",
2669
  "modality": "text_to_video",
2670
  "data": text_to_video_df,
2671
  "columns": text_to_video_display_columns,
 
2665
  datasets = [
2666
  {
2667
  "id": "text_to_video",
2668
+ "name": "VBench-2.0 Dataset",
2669
  "modality": "text_to_video",
2670
  "data": text_to_video_df,
2671
  "columns": text_to_video_display_columns,
ui.py CHANGED
@@ -80,8 +80,8 @@ and price**. Each view is a **dataset** scored with a **metric**, written as
80
 
81
  ## How to read it
82
 
83
- 1. Pick a **type** (Text to Video, Video to Video, or Text to Image), then a
84
- **dataset** and a **metric**.
85
  2. **Leaderboards**: ranked by that metric. Price and generation time sit in
86
  the same table when the source publishes them.
87
  3. **Pareto plots**: mark models that are not beaten on both higher score
@@ -109,12 +109,12 @@ prompt suites, so samples are not shown.
109
 
110
  ## Current datasets
111
 
112
- ### Pruna Internal Text-to-Video Benchmark
113
- Pruna's internal text-to-video benchmark, comparing P-Video-2 variants with
114
- Fal-hosted models. Quality is Datapoint Elo and Rapidata Elo from pairwise
115
- preference. Price is USD per second of output video. Time per second of
116
- video is Fal wall time, except Pruna models which use model execution time.
117
- Samples are not shown for this dataset.
118
 
119
  ### Pruna Internal Video-Edit Benchmark
120
  Pruna's internal video-to-video editing benchmark, collected by our
@@ -193,7 +193,7 @@ models *within* a Dataset | Metric view.
193
  - **Prompt counts:** OneIG Alignment uses 100 anime, 100 human, and 99 object
194
  prompts (299 total). Qwen Image Dataset uses 100 prompts sampled from the
195
  1,000-prompt pool for roughly even coverage of its fine-grained (L3)
196
- categories. The Pruna Internal Text-to-Video Benchmark uses about 90
197
  generations per model. The Pruna Internal Video-Edit Benchmark uses 78
198
  prompts across advertising, e-commerce, real estate, camera, lighting,
199
  text, and related categories. Artificial Analysis and Arena AI use their
@@ -1385,11 +1385,11 @@ def _filter_row(datasets, metrics, default_dataset_id, default_metric_id=None):
1385
  modality_dd = gr.Dropdown(
1386
  choices=_modality_choices(datasets),
1387
  value=default_modality,
1388
- label="Type",
1389
  type="value",
1390
  filterable=False,
1391
  scale=1,
1392
- min_width=150,
1393
  )
1394
  dataset_dd = gr.Dropdown(
1395
  choices=_dataset_choices(datasets, modality=default_modality),
@@ -1439,12 +1439,13 @@ def render_image_workspace(datasets, metrics, default_dataset_id, default_metric
1439
  with gr.Column(elem_classes="workspace-filters") as filters_host:
1440
  gr.Markdown(
1441
  "<p class='filter-help'>"
1442
- "Start with Type to switch between Video to Video and Text "
1443
- "to Image. The rest of the filters follow you across "
1444
- "Leaderboards, Pareto plots, and Samples. Samples only "
1445
- "lists datasets and models we have generations for; Pareto "
1446
- "plots only lists datasets with price or generation time. "
1447
- "Search in Models, or leave it empty to include every model."
 
1448
  "</p>",
1449
  elem_classes="filter-help-host",
1450
  )
 
80
 
81
  ## How to read it
82
 
83
+ 1. Pick a **model type** (Text to Video, Video to Video, or Text to Image),
84
+ then a **dataset** and a **metric**.
85
  2. **Leaderboards**: ranked by that metric. Price and generation time sit in
86
  the same table when the source publishes them.
87
  3. **Pareto plots**: mark models that are not beaten on both higher score
 
109
 
110
  ## Current datasets
111
 
112
+ ### VBench-2.0 Dataset
113
+ VBench-2.0 prompts, comparing P-Video-2 variants with Fal-hosted models.
114
+ Quality is Datapoint Elo and Rapidata Elo from pairwise preference. Price
115
+ is USD per second of output video. Time per second of video is Fal wall
116
+ time, except Pruna models which use model execution time. Samples are not
117
+ shown for this dataset.
118
 
119
  ### Pruna Internal Video-Edit Benchmark
120
  Pruna's internal video-to-video editing benchmark, collected by our
 
193
  - **Prompt counts:** OneIG Alignment uses 100 anime, 100 human, and 99 object
194
  prompts (299 total). Qwen Image Dataset uses 100 prompts sampled from the
195
  1,000-prompt pool for roughly even coverage of its fine-grained (L3)
196
+ categories. The VBench-2.0 Dataset uses about 90
197
  generations per model. The Pruna Internal Video-Edit Benchmark uses 78
198
  prompts across advertising, e-commerce, real estate, camera, lighting,
199
  text, and related categories. Artificial Analysis and Arena AI use their
 
1385
  modality_dd = gr.Dropdown(
1386
  choices=_modality_choices(datasets),
1387
  value=default_modality,
1388
+ label="Model Type",
1389
  type="value",
1390
  filterable=False,
1391
  scale=1,
1392
+ min_width=170,
1393
  )
1394
  dataset_dd = gr.Dropdown(
1395
  choices=_dataset_choices(datasets, modality=default_modality),
 
1439
  with gr.Column(elem_classes="workspace-filters") as filters_host:
1440
  gr.Markdown(
1441
  "<p class='filter-help'>"
1442
+ "Start with Model Type to switch modalities. The rest of "
1443
+ "the filters follow you across Leaderboards, Pareto plots, "
1444
+ "and Samples. "
1445
+ "Samples only lists datasets and models we have generations "
1446
+ "for; Pareto plots only lists datasets with price or "
1447
+ "generation time. Search in Models, or leave it empty to "
1448
+ "include every model."
1449
  "</p>",
1450
  elem_classes="filter-help-host",
1451
  )