DecisionLab / app /static /index.html
Michael Stattelman
Add application file
6012dcc
Raw History Blame Contribute Delete
11.5 kB
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>DecisionLab · your decision models vs Laya</title>
<meta name="description" content="Put your local decision models (LightDec, Arthur) and Laya on the same decisions and compare their answers, confidence and speed. Runs locally.">
<link rel="icon" href="/static/logo.png">
<link rel="stylesheet" href="/static/app.css">
</head>
<body>
<header class="topbar">
<a class="brand" href="#top" aria-label="DecisionLab home">
<img src="/static/logo.png" alt="Falcons.ai" class="brand-logo">
<span class="brand-name">DecisionLab</span>
</a>
<nav class="nav" aria-label="Sections">
<a href="#setup">Setup</a>
<a href="#decision">Decision</a>
<a href="#verdicts">Verdicts</a>
<a href="#scoreboard">Scoreboard</a>
</nav>
<div class="readiness" id="readiness" aria-live="polite">Loading models…</div>
</header>
<main id="top">
<section class="hero">
<div class="hero-copy">
<p class="kicker">Your LightDec models and Laya, on your own decisions</p>
<h1>Your decision models.<br>The same question.</h1>
<p class="lede">Every model reads a state, answer typed questions and return a probability for every option,
without writing a word. Give them the same decision and see where they agree, how sure they are, and how fast they answer.</p>
<div class="hero-actions">
<a class="btn btn-primary" href="#decision">Try a demo</a>
<a class="btn btn-quiet" href="#scoreboard">Run the scoreboard</a>
</div>
</div>
<div class="versus" id="versus" aria-hidden="true"></div>
</section>
<section id="setup" class="section">
<header class="section-head">
<p class="step">00 / setup</p>
<h2 data-n-models="{N} models load when the lab starts">Models load when the lab starts</h2>
<p class="section-note">The first start downloads LightDec_Arthur, LightDec_V2, Enterprise Reflux Laya V2.1 and Laya from Hugging Face (about 2 GB) into the cache volume. Every LightDec (FalconDec) and Arthur model in the local <code>models</code> folder is found at start-up and loaded from there. Nothing you type leaves this server.</p>
<p class="section-note" id="models-found"></p>
</header>
<div class="models" id="models">
</div>
<p class="env" id="env"></p>
</section>
<section id="decision" class="section">
<header class="section-head">
<p class="step">01 / decision</p>
<h2>Give them a real choice</h2>
<p class="section-note">Pick a demo or write your own. Questions use the same JSON shape as the Laya and Jev SDKs:
<code>choice</code> for labelled options, <code>score</code> for an ordered scale, <code>noul</code> for yes or no.</p>
</header>
<div class="picker" aria-label="Demos">
<div class="picker-bar">
<div class="tabs-wrap" id="tabs" role="tablist" aria-label="Demo groups"></div>
<label class="search">
<span class="visually-hidden">Filter demos</span>
<input type="search" id="demo-filter" placeholder="Filter all demos" autocomplete="off">
</label>
</div>
<p class="group-blurb" id="group-blurb"></p>
<div class="tiles" id="tiles" role="listbox" aria-label="Demos"></div>
</div>
<div class="current">
<div class="current-text">
<p class="current-title" id="demo-title">Choose a demo</p>
<p class="demo-blurb" id="demo-blurb"></p>
</div>
<div class="current-nav">
<button type="button" class="btn btn-quiet btn-small" id="prev" aria-label="Previous demo">Previous</button>
<button type="button" class="btn btn-quiet btn-small" id="next" aria-label="Next demo">Next</button>
</div>
</div>
<div class="editors">
<label class="field">
<span>State <small>text or JSON</small></span>
<textarea id="state" rows="15" spellcheck="false"></textarea>
</label>
<label class="field">
<span>Questions <small>JSON</small></span>
<textarea id="questions" rows="15" spellcheck="false"></textarea>
</label>
</div>
<div class="controls">
<label class="control">
<span>Act automatically when confidence ≥ <output id="thr-out">0.80</output></span>
<input type="range" id="threshold" min="0.50" max="0.99" step="0.01" value="0.80">
</label>
<label class="control">
<span>Confidence means</span>
<select id="measure">
<option value="top_prob">Probability of the top option (LightDec, Jev)</option>
<option value="laya_conf">Laya's definition (1 − normalised entropy; top probability for yes/no)</option>
</select>
</label>
<button class="btn btn-primary run" id="run" type="button">Run the models</button>
</div>
<p class="form-error" id="form-error" role="alert"></p>
</section>
<section id="verdicts" class="section">
<header class="section-head">
<p class="step">02 / verdicts</p>
<h2>Side by side</h2>
<p class="section-note">Each option shows one bar per model, in the order of the cards above; the bold bar is that model's top answer. A verdict is marked <em>act</em> when its confidence clears your threshold and <em>defer</em> when a person should decide.</p>
</header>
<div class="timing" id="timing" hidden></div>
<div class="verdicts" id="verdict-list">
<p class="empty">Run a decision to see every model's answer here.</p>
</div>
<details class="raw" id="raw" hidden>
<summary>Raw JSON</summary>
<pre id="raw-json"></pre>
</details>
</section>
<section id="scoreboard" class="section">
<header class="section-head">
<p class="step">03 / scoreboard</p>
<h2>Every demo, ranked for agents</h2>
<p class="section-note">Runs every demo through every model, compares each answer with a reference answer (a careful human reading, not ground truth), highlights the better model on each metric and ranks them by the Agentic Use Score. A few dozen demos show behaviour, not accuracy; test on your own labelled data before trusting a threshold.</p>
</header>
<div class="score-actions">
<label class="scope">
<span>Run</span>
<select id="scope" aria-label="Which demos to run"><option value="">All demos</option></select>
</label>
<button class="btn btn-primary" id="run-all" type="button">Run all demos</button>
<span class="progress" id="progress" aria-live="polite"></span>
</div>
<div class="summary" id="summary" hidden></div>
<div class="table-wrap">
<table class="score-table" id="score-table" hidden>
<thead>
<tr id="score-head"></tr>
</thead>
<tbody></tbody>
</table>
</div>
</section>
<section class="section notes">
<header class="section-head">
<p class="step">04 / reading the numbers</p>
<h2>What these numbers do, and do not, mean</h2>
</header>
<div class="aus-guide">
<div class="aus-intro">
<h3>The Agentic Use Score</h3>
<p>When an agent acts on an answer, the mistakes that hurt are the confident ones: the table it deletes, the payment it sends.
A mistake the model defers costs only a quick human look. A correct answer it defers costs a little time.</p>
<p>So the score, from 0 to 100, rewards a model most for being right when it acts and for flagging its own mistakes.</p>
<p class="aus-footnote">High-stakes questions (irreversible actions, security, money, personal data) count twice, except in speed.
Moving the threshold or changing the confidence measure changes the score, because together they decide when a model acts.</p>
</div>
<ol class="weights">
<li><span class="w-pct">35%</span><span class="w-bar"><i class="w-100"></i></span>
<span class="w-text"><b>Right when it acts</b>Of the answers confident enough to act on, how many were correct. A small allowance stops a model that acts only once or twice from scoring perfectly.</span></li>
<li><span class="w-pct">25%</span><span class="w-bar"><i class="w-71"></i></span>
<span class="w-text"><b>Flags its own mistakes</b>Of the answers that were wrong, how many were unsure enough to go to a person instead.</span></li>
<li><span class="w-pct">15%</span><span class="w-bar"><i class="w-43"></i></span>
<span class="w-text"><b>Overall accuracy</b>How many answers match the reference, acted on or not.</span></li>
<li><span class="w-pct">15%</span><span class="w-bar"><i class="w-43"></i></span>
<span class="w-text"><b>Handles on its own</b>How many decisions clear the threshold, so the agent doesn't need a person.</span></li>
<li><span class="w-pct">10%</span><span class="w-bar"><i class="w-29"></i></span>
<span class="w-text"><b>Speed</b>Full marks at 50 ms or faster, none at one second or slower. Agents make many decisions per task.</span></li>
</ol>
</div>
<div class="note-grid">
<div class="note">
<h3>Two kinds of confidence</h3>
<p>LightDec (both versions) and Jev report the probability of the top option.</p>
<p>Laya reports how concentrated the whole distribution is (1 − normalised entropy) for choice and score questions, and the larger of P(yes) and P(no) for yes/no.</p>
<p>The lab computes both measures for every model. Pick one with <em>Confidence means</em>.</p>
</div>
<div class="note">
<h3>Real local timing</h3>
<p>Each model is timed on this server with a high-resolution clock.</p>
<p>The models run one after the other, so they never compete for the device.</p>
<p>Times depend on the hardware: a GPU is much faster than a CPU.</p>
</div>
<div class="note">
<h3>The models</h3>
<p><a href="https://huggingface.co/Falconsai/LightDec_Arthur" target="_blank" rel="noopener noreferrer">LightDec_Arthur</a>: byte-level Arthur model, about 12M parameters; the code that runs it ships with DecisionLab.</p>
<p><a href="https://huggingface.co/Falconsai/LightDec_V2" target="_blank" rel="noopener noreferrer">LightDec_V2</a>: FalconDec on the Ettin-150M encoder, 2,048-token window, Apache-2.0.</p>
<p><a href="https://huggingface.co/yasserrmd/enterprise-reflux-laya-v21" target="_blank" rel="noopener noreferrer">Enterprise Reflux Laya V2.1</a>: a fine-tune of Laya for ranking enterprise actions, by Mohamed Yasser; its model card marks it a research prototype, not for production.</p>
<p><a href="https://huggingface.co/convaiinnovations/laya" target="_blank" rel="noopener noreferrer">Laya</a>: by Convai Innovations, ModernBERT-large, Apache-2.0.</p>
<p>Models in the mounted <code>models</code> folder are added after these, marked "(local)": a LightDec folder has a <code>falcondec_config.json</code>, an Arthur folder a <code>config.json</code> and <code>model.safetensors</code>.</p>
<p>None of the models writes text. Each only ranks the options you give it.</p>
</div>
</div>
</section>
</main>
<footer class="footer">
<img src="/static/logo.png" alt="Falcons.ai" class="footer-logo">
<span>DecisionLab</span>
</footer>
<script src="/static/app.js"></script>
</body>
</html>