Spaces:
Running on Zero
Running on Zero
File size: 11,493 Bytes
66ee87e 6012dcc 66ee87e 6012dcc 66ee87e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 | <!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>DecisionLab · your decision models vs Laya</title>
<meta name="description" content="Put your local decision models (LightDec, Arthur) and Laya on the same decisions and compare their answers, confidence and speed. Runs locally.">
<link rel="icon" href="/static/logo.png">
<link rel="stylesheet" href="/static/app.css">
</head>
<body>
<header class="topbar">
<a class="brand" href="#top" aria-label="DecisionLab home">
<img src="/static/logo.png" alt="Falcons.ai" class="brand-logo">
<span class="brand-name">DecisionLab</span>
</a>
<nav class="nav" aria-label="Sections">
<a href="#setup">Setup</a>
<a href="#decision">Decision</a>
<a href="#verdicts">Verdicts</a>
<a href="#scoreboard">Scoreboard</a>
</nav>
<div class="readiness" id="readiness" aria-live="polite">Loading models…</div>
</header>
<main id="top">
<section class="hero">
<div class="hero-copy">
<p class="kicker">Your LightDec models and Laya, on your own decisions</p>
<h1>Your decision models.<br>The same question.</h1>
<p class="lede">Every model reads a state, answer typed questions and return a probability for every option,
without writing a word. Give them the same decision and see where they agree, how sure they are, and how fast they answer.</p>
<div class="hero-actions">
<a class="btn btn-primary" href="#decision">Try a demo</a>
<a class="btn btn-quiet" href="#scoreboard">Run the scoreboard</a>
</div>
</div>
<div class="versus" id="versus" aria-hidden="true"></div>
</section>
<section id="setup" class="section">
<header class="section-head">
<p class="step">00 / setup</p>
<h2 data-n-models="{N} models load when the lab starts">Models load when the lab starts</h2>
<p class="section-note">The first start downloads LightDec_Arthur, LightDec_V2, Enterprise Reflux Laya V2.1 and Laya from Hugging Face (about 2 GB) into the cache volume. Every LightDec (FalconDec) and Arthur model in the local <code>models</code> folder is found at start-up and loaded from there. Nothing you type leaves this server.</p>
<p class="section-note" id="models-found"></p>
</header>
<div class="models" id="models">
</div>
<p class="env" id="env"></p>
</section>
<section id="decision" class="section">
<header class="section-head">
<p class="step">01 / decision</p>
<h2>Give them a real choice</h2>
<p class="section-note">Pick a demo or write your own. Questions use the same JSON shape as the Laya and Jev SDKs:
<code>choice</code> for labelled options, <code>score</code> for an ordered scale, <code>noul</code> for yes or no.</p>
</header>
<div class="picker" aria-label="Demos">
<div class="picker-bar">
<div class="tabs-wrap" id="tabs" role="tablist" aria-label="Demo groups"></div>
<label class="search">
<span class="visually-hidden">Filter demos</span>
<input type="search" id="demo-filter" placeholder="Filter all demos" autocomplete="off">
</label>
</div>
<p class="group-blurb" id="group-blurb"></p>
<div class="tiles" id="tiles" role="listbox" aria-label="Demos"></div>
</div>
<div class="current">
<div class="current-text">
<p class="current-title" id="demo-title">Choose a demo</p>
<p class="demo-blurb" id="demo-blurb"></p>
</div>
<div class="current-nav">
<button type="button" class="btn btn-quiet btn-small" id="prev" aria-label="Previous demo">Previous</button>
<button type="button" class="btn btn-quiet btn-small" id="next" aria-label="Next demo">Next</button>
</div>
</div>
<div class="editors">
<label class="field">
<span>State <small>text or JSON</small></span>
<textarea id="state" rows="15" spellcheck="false"></textarea>
</label>
<label class="field">
<span>Questions <small>JSON</small></span>
<textarea id="questions" rows="15" spellcheck="false"></textarea>
</label>
</div>
<div class="controls">
<label class="control">
<span>Act automatically when confidence ≥ <output id="thr-out">0.80</output></span>
<input type="range" id="threshold" min="0.50" max="0.99" step="0.01" value="0.80">
</label>
<label class="control">
<span>Confidence means</span>
<select id="measure">
<option value="top_prob">Probability of the top option (LightDec, Jev)</option>
<option value="laya_conf">Laya's definition (1 − normalised entropy; top probability for yes/no)</option>
</select>
</label>
<button class="btn btn-primary run" id="run" type="button">Run the models</button>
</div>
<p class="form-error" id="form-error" role="alert"></p>
</section>
<section id="verdicts" class="section">
<header class="section-head">
<p class="step">02 / verdicts</p>
<h2>Side by side</h2>
<p class="section-note">Each option shows one bar per model, in the order of the cards above; the bold bar is that model's top answer. A verdict is marked <em>act</em> when its confidence clears your threshold and <em>defer</em> when a person should decide.</p>
</header>
<div class="timing" id="timing" hidden></div>
<div class="verdicts" id="verdict-list">
<p class="empty">Run a decision to see every model's answer here.</p>
</div>
<details class="raw" id="raw" hidden>
<summary>Raw JSON</summary>
<pre id="raw-json"></pre>
</details>
</section>
<section id="scoreboard" class="section">
<header class="section-head">
<p class="step">03 / scoreboard</p>
<h2>Every demo, ranked for agents</h2>
<p class="section-note">Runs every demo through every model, compares each answer with a reference answer (a careful human reading, not ground truth), highlights the better model on each metric and ranks them by the Agentic Use Score. A few dozen demos show behaviour, not accuracy; test on your own labelled data before trusting a threshold.</p>
</header>
<div class="score-actions">
<label class="scope">
<span>Run</span>
<select id="scope" aria-label="Which demos to run"><option value="">All demos</option></select>
</label>
<button class="btn btn-primary" id="run-all" type="button">Run all demos</button>
<span class="progress" id="progress" aria-live="polite"></span>
</div>
<div class="summary" id="summary" hidden></div>
<div class="table-wrap">
<table class="score-table" id="score-table" hidden>
<thead>
<tr id="score-head"></tr>
</thead>
<tbody></tbody>
</table>
</div>
</section>
<section class="section notes">
<header class="section-head">
<p class="step">04 / reading the numbers</p>
<h2>What these numbers do, and do not, mean</h2>
</header>
<div class="aus-guide">
<div class="aus-intro">
<h3>The Agentic Use Score</h3>
<p>When an agent acts on an answer, the mistakes that hurt are the confident ones: the table it deletes, the payment it sends.
A mistake the model defers costs only a quick human look. A correct answer it defers costs a little time.</p>
<p>So the score, from 0 to 100, rewards a model most for being right when it acts and for flagging its own mistakes.</p>
<p class="aus-footnote">High-stakes questions (irreversible actions, security, money, personal data) count twice, except in speed.
Moving the threshold or changing the confidence measure changes the score, because together they decide when a model acts.</p>
</div>
<ol class="weights">
<li><span class="w-pct">35%</span><span class="w-bar"><i class="w-100"></i></span>
<span class="w-text"><b>Right when it acts</b>Of the answers confident enough to act on, how many were correct. A small allowance stops a model that acts only once or twice from scoring perfectly.</span></li>
<li><span class="w-pct">25%</span><span class="w-bar"><i class="w-71"></i></span>
<span class="w-text"><b>Flags its own mistakes</b>Of the answers that were wrong, how many were unsure enough to go to a person instead.</span></li>
<li><span class="w-pct">15%</span><span class="w-bar"><i class="w-43"></i></span>
<span class="w-text"><b>Overall accuracy</b>How many answers match the reference, acted on or not.</span></li>
<li><span class="w-pct">15%</span><span class="w-bar"><i class="w-43"></i></span>
<span class="w-text"><b>Handles on its own</b>How many decisions clear the threshold, so the agent doesn't need a person.</span></li>
<li><span class="w-pct">10%</span><span class="w-bar"><i class="w-29"></i></span>
<span class="w-text"><b>Speed</b>Full marks at 50 ms or faster, none at one second or slower. Agents make many decisions per task.</span></li>
</ol>
</div>
<div class="note-grid">
<div class="note">
<h3>Two kinds of confidence</h3>
<p>LightDec (both versions) and Jev report the probability of the top option.</p>
<p>Laya reports how concentrated the whole distribution is (1 − normalised entropy) for choice and score questions, and the larger of P(yes) and P(no) for yes/no.</p>
<p>The lab computes both measures for every model. Pick one with <em>Confidence means</em>.</p>
</div>
<div class="note">
<h3>Real local timing</h3>
<p>Each model is timed on this server with a high-resolution clock.</p>
<p>The models run one after the other, so they never compete for the device.</p>
<p>Times depend on the hardware: a GPU is much faster than a CPU.</p>
</div>
<div class="note">
<h3>The models</h3>
<p><a href="https://huggingface.co/Falconsai/LightDec_Arthur" target="_blank" rel="noopener noreferrer">LightDec_Arthur</a>: byte-level Arthur model, about 12M parameters; the code that runs it ships with DecisionLab.</p>
<p><a href="https://huggingface.co/Falconsai/LightDec_V2" target="_blank" rel="noopener noreferrer">LightDec_V2</a>: FalconDec on the Ettin-150M encoder, 2,048-token window, Apache-2.0.</p>
<p><a href="https://huggingface.co/yasserrmd/enterprise-reflux-laya-v21" target="_blank" rel="noopener noreferrer">Enterprise Reflux Laya V2.1</a>: a fine-tune of Laya for ranking enterprise actions, by Mohamed Yasser; its model card marks it a research prototype, not for production.</p>
<p><a href="https://huggingface.co/convaiinnovations/laya" target="_blank" rel="noopener noreferrer">Laya</a>: by Convai Innovations, ModernBERT-large, Apache-2.0.</p>
<p>Models in the mounted <code>models</code> folder are added after these, marked "(local)": a LightDec folder has a <code>falcondec_config.json</code>, an Arthur folder a <code>config.json</code> and <code>model.safetensors</code>.</p>
<p>None of the models writes text. Each only ranks the options you give it.</p>
</div>
</div>
</section>
</main>
<footer class="footer">
<img src="/static/logo.png" alt="Falcons.ai" class="footer-logo">
<span>DecisionLab</span>
</footer>
<script src="/static/app.js"></script>
</body>
</html>
|