Download README.md from FabcanK6/ivis-bert: direct link, hf CLI and curl.
- Browser
- Download file 4.65 kB
-
https://huggingface.co/FabcanK6/ivis-bert/resolve/main/README.md
- Command line
-
hf download hf://FabcanK6/ivis-bert/README.md
-
curl -L -o README.md https://huggingface.co/FabcanK6/ivis-bert/resolve/main/README.md
license: mit
language:
- en
base_model: bert-base-uncased
library_name: pytorch
tags:
- power-bi
- text-to-visualization
- intent-classification
- slot-filling
- joint-bert
- clinical-trials
- clinical-operations
iVIS BERT: plain-English dashboard requests → Power BI visuals
Try it live: ivis-fabcank6.streamlit.app · Code: github.com/FabcanK6/ivis
This is the model behind iVIS (Intelligent Visualization Insight Synthesizer). It reads a loose request such as "Can we get a chart of the top 10 sites with the most open queries in Germany over the last 30 days?" and finds what was asked for: the chart type and each part of the request (measure, axis, legend, filters, time window, top N, sort). iVIS then maps those parts onto a Power BI data model and returns a visual spec, a Power BI visual file (PBIR visual.json) and a work item.
"top 10 sites with the most open queries in Germany last 30 days"
chart: bar → clusteredBarChart measure: Queries[Query Count]
axis: Site[Site Name] filters: Queries[Status] = open, Site[Country] = Germany
top N: 10, descending date: last 30 days
What the model does
A fine-tuned bert-base-uncased encoder with two heads, trained jointly (the "JointBERT" set-up):
| Head | Output |
|---|---|
Chart type (from the [CLS] vector) |
one of 11: bar, stacked_bar, line, area, pie, donut, table, matrix, card, scatter, map |
| Slots (one tag per word, BIO) | METRIC, AGG, GROUP_BY, SERIES, FILTER, TIME, TOPN, SORT |
The model only decides what was asked for. Which Power BI fields that means comes from a data model (catalog) in the iVIS code, so the same model can be pointed at another dataset's measures and columns.
How to use it
The model uses a custom head, so load it with the iVIS code rather than a transformers pipeline:
git clone https://github.com/FabcanK6/ivis && cd ivis
pip install -r requirements.txt
from huggingface_hub import snapshot_download
from ivis.predict import BertParser, HybridParser
path = snapshot_download("FabcanK6/ivis-bert")
parser = HybridParser(BertParser(path)) # BERT + strong keyword cues (recommended)
spec = parser.parse("Protocol deviations by country broken down by deviation category in 2025")
print(spec["powerbi_visual"], spec["field_wells"])
Or from the command line: python -m ivis.cli --format visual "share of SAEs by country this year" > visual.json
Results
2,000 test requests, half using phrasings the model never saw in training (10 held-out templates).
| Reader | Chart type right | Same visual | New wording: chart right | New wording: same visual |
|---|---|---|---|---|
| Keywords only (baseline) | 94.5% | 73.0% | 95.6% | 76.0% |
| This model alone | 92.6% | 83.4% | 85.3% | 66.8% |
| This model + keyword cues (default in the app) | 99.6% | 85.1% | 99.4% | 70.5% |
Same visual means the Power BI visual built from the reading matches the one built from the labels (chart, measures, axis, legend, filters, time window, top N). The model's slot F1 is 0.962 and its exact frame match 83.0%. When a request names the chart's layout ("split by", "headline", "donut"), the app lets those words decide the chart, which closes the model's gap on new wording.
Training
- Data: 20,000 synthetic requests from a template generator (about 85 templates across the 11 chart types, deliberately indirect phrasing, conversational prefixes, casing changes and typos), with slots filled from a clinical-operations data model: 14 measures (queries, SAEs, deviations, enrollment, SDV, …) and 20 dimensions (site, country, study, visit, arm, …). The generator is in the repo (
ivis/data/generate.py). - Setup:
bert-base-uncased, 4 epochs, AdamW with warmup, on a free Colab T4.
Limitations
- Trained on generated requests. Real requests are messier, so expect lower accuracy than the table above. Evaluating on hand-labelled real requests is the next step.
- Learned one data model's wording. With a very different dataset, the model still finds the parts of a request, but the app's AI reader (which reads your own field list) usually maps them better.
- New phrasing. On unfamiliar wording the model sometimes mixes up the axis and the legend; that is the main remaining gap.
- English only.
About
Built by Fabian Msafiri. All data used to train and test the model is synthetic; no real study or patient data was used. Licensed MIT.