Spaces:
Running on Zero
Running on Zero
Download app/static/index.html from RealFalconsAI/DecisionLab: direct link, hf CLI and curl.
- Browser
- Download file 11.5 kB
-
https://huggingface.co/spaces/RealFalconsAI/DecisionLab/resolve/main/app/static/index.html
- Command line
-
hf download hf://spaces/RealFalconsAI/DecisionLab/app/static/index.html
-
curl -L -o index.html https://huggingface.co/spaces/RealFalconsAI/DecisionLab/resolve/main/app/static/index.html
11.5 kB
| <html lang="en"> | |
| <head> | |
| <meta charset="utf-8"> | |
| <meta name="viewport" content="width=device-width, initial-scale=1"> | |
| <title>DecisionLab · your decision models vs Laya</title> | |
| <meta name="description" content="Put your local decision models (LightDec, Arthur) and Laya on the same decisions and compare their answers, confidence and speed. Runs locally."> | |
| <link rel="icon" href="/static/logo.png"> | |
| <link rel="stylesheet" href="/static/app.css"> | |
| </head> | |
| <body> | |
| <header class="topbar"> | |
| <a class="brand" href="#top" aria-label="DecisionLab home"> | |
| <img src="/static/logo.png" alt="Falcons.ai" class="brand-logo"> | |
| <span class="brand-name">DecisionLab</span> | |
| </a> | |
| <nav class="nav" aria-label="Sections"> | |
| <a href="#setup">Setup</a> | |
| <a href="#decision">Decision</a> | |
| <a href="#verdicts">Verdicts</a> | |
| <a href="#scoreboard">Scoreboard</a> | |
| </nav> | |
| <div class="readiness" id="readiness" aria-live="polite">Loading models…</div> | |
| </header> | |
| <main id="top"> | |
| <section class="hero"> | |
| <div class="hero-copy"> | |
| <p class="kicker">Your LightDec models and Laya, on your own decisions</p> | |
| <h1>Your decision models.<br>The same question.</h1> | |
| <p class="lede">Every model reads a state, answer typed questions and return a probability for every option, | |
| without writing a word. Give them the same decision and see where they agree, how sure they are, and how fast they answer.</p> | |
| <div class="hero-actions"> | |
| <a class="btn btn-primary" href="#decision">Try a demo</a> | |
| <a class="btn btn-quiet" href="#scoreboard">Run the scoreboard</a> | |
| </div> | |
| </div> | |
| <div class="versus" id="versus" aria-hidden="true"></div> | |
| </section> | |
| <section id="setup" class="section"> | |
| <header class="section-head"> | |
| <p class="step">00 / setup</p> | |
| <h2 data-n-models="{N} models load when the lab starts">Models load when the lab starts</h2> | |
| <p class="section-note">The first start downloads LightDec_Arthur, LightDec_V2, Enterprise Reflux Laya V2.1 and Laya from Hugging Face (about 2 GB) into the cache volume. Every LightDec (FalconDec) and Arthur model in the local <code>models</code> folder is found at start-up and loaded from there. Nothing you type leaves this server.</p> | |
| <p class="section-note" id="models-found"></p> | |
| </header> | |
| <div class="models" id="models"> | |
| </div> | |
| <p class="env" id="env"></p> | |
| </section> | |
| <section id="decision" class="section"> | |
| <header class="section-head"> | |
| <p class="step">01 / decision</p> | |
| <h2>Give them a real choice</h2> | |
| <p class="section-note">Pick a demo or write your own. Questions use the same JSON shape as the Laya and Jev SDKs: | |
| <code>choice</code> for labelled options, <code>score</code> for an ordered scale, <code>noul</code> for yes or no.</p> | |
| </header> | |
| <div class="picker" aria-label="Demos"> | |
| <div class="picker-bar"> | |
| <div class="tabs-wrap" id="tabs" role="tablist" aria-label="Demo groups"></div> | |
| <label class="search"> | |
| <span class="visually-hidden">Filter demos</span> | |
| <input type="search" id="demo-filter" placeholder="Filter all demos" autocomplete="off"> | |
| </label> | |
| </div> | |
| <p class="group-blurb" id="group-blurb"></p> | |
| <div class="tiles" id="tiles" role="listbox" aria-label="Demos"></div> | |
| </div> | |
| <div class="current"> | |
| <div class="current-text"> | |
| <p class="current-title" id="demo-title">Choose a demo</p> | |
| <p class="demo-blurb" id="demo-blurb"></p> | |
| </div> | |
| <div class="current-nav"> | |
| <button type="button" class="btn btn-quiet btn-small" id="prev" aria-label="Previous demo">Previous</button> | |
| <button type="button" class="btn btn-quiet btn-small" id="next" aria-label="Next demo">Next</button> | |
| </div> | |
| </div> | |
| <div class="editors"> | |
| <label class="field"> | |
| <span>State <small>text or JSON</small></span> | |
| <textarea id="state" rows="15" spellcheck="false"></textarea> | |
| </label> | |
| <label class="field"> | |
| <span>Questions <small>JSON</small></span> | |
| <textarea id="questions" rows="15" spellcheck="false"></textarea> | |
| </label> | |
| </div> | |
| <div class="controls"> | |
| <label class="control"> | |
| <span>Act automatically when confidence ≥ <output id="thr-out">0.80</output></span> | |
| <input type="range" id="threshold" min="0.50" max="0.99" step="0.01" value="0.80"> | |
| </label> | |
| <label class="control"> | |
| <span>Confidence means</span> | |
| <select id="measure"> | |
| <option value="top_prob">Probability of the top option (LightDec, Jev)</option> | |
| <option value="laya_conf">Laya's definition (1 − normalised entropy; top probability for yes/no)</option> | |
| </select> | |
| </label> | |
| <button class="btn btn-primary run" id="run" type="button">Run the models</button> | |
| </div> | |
| <p class="form-error" id="form-error" role="alert"></p> | |
| </section> | |
| <section id="verdicts" class="section"> | |
| <header class="section-head"> | |
| <p class="step">02 / verdicts</p> | |
| <h2>Side by side</h2> | |
| <p class="section-note">Each option shows one bar per model, in the order of the cards above; the bold bar is that model's top answer. A verdict is marked <em>act</em> when its confidence clears your threshold and <em>defer</em> when a person should decide.</p> | |
| </header> | |
| <div class="timing" id="timing" hidden></div> | |
| <div class="verdicts" id="verdict-list"> | |
| <p class="empty">Run a decision to see every model's answer here.</p> | |
| </div> | |
| <details class="raw" id="raw" hidden> | |
| <summary>Raw JSON</summary> | |
| <pre id="raw-json"></pre> | |
| </details> | |
| </section> | |
| <section id="scoreboard" class="section"> | |
| <header class="section-head"> | |
| <p class="step">03 / scoreboard</p> | |
| <h2>Every demo, ranked for agents</h2> | |
| <p class="section-note">Runs every demo through every model, compares each answer with a reference answer (a careful human reading, not ground truth), highlights the better model on each metric and ranks them by the Agentic Use Score. A few dozen demos show behaviour, not accuracy; test on your own labelled data before trusting a threshold.</p> | |
| </header> | |
| <div class="score-actions"> | |
| <label class="scope"> | |
| <span>Run</span> | |
| <select id="scope" aria-label="Which demos to run"><option value="">All demos</option></select> | |
| </label> | |
| <button class="btn btn-primary" id="run-all" type="button">Run all demos</button> | |
| <span class="progress" id="progress" aria-live="polite"></span> | |
| </div> | |
| <div class="summary" id="summary" hidden></div> | |
| <div class="table-wrap"> | |
| <table class="score-table" id="score-table" hidden> | |
| <thead> | |
| <tr id="score-head"></tr> | |
| </thead> | |
| <tbody></tbody> | |
| </table> | |
| </div> | |
| </section> | |
| <section class="section notes"> | |
| <header class="section-head"> | |
| <p class="step">04 / reading the numbers</p> | |
| <h2>What these numbers do, and do not, mean</h2> | |
| </header> | |
| <div class="aus-guide"> | |
| <div class="aus-intro"> | |
| <h3>The Agentic Use Score</h3> | |
| <p>When an agent acts on an answer, the mistakes that hurt are the confident ones: the table it deletes, the payment it sends. | |
| A mistake the model defers costs only a quick human look. A correct answer it defers costs a little time.</p> | |
| <p>So the score, from 0 to 100, rewards a model most for being right when it acts and for flagging its own mistakes.</p> | |
| <p class="aus-footnote">High-stakes questions (irreversible actions, security, money, personal data) count twice, except in speed. | |
| Moving the threshold or changing the confidence measure changes the score, because together they decide when a model acts.</p> | |
| </div> | |
| <ol class="weights"> | |
| <li><span class="w-pct">35%</span><span class="w-bar"><i class="w-100"></i></span> | |
| <span class="w-text"><b>Right when it acts</b>Of the answers confident enough to act on, how many were correct. A small allowance stops a model that acts only once or twice from scoring perfectly.</span></li> | |
| <li><span class="w-pct">25%</span><span class="w-bar"><i class="w-71"></i></span> | |
| <span class="w-text"><b>Flags its own mistakes</b>Of the answers that were wrong, how many were unsure enough to go to a person instead.</span></li> | |
| <li><span class="w-pct">15%</span><span class="w-bar"><i class="w-43"></i></span> | |
| <span class="w-text"><b>Overall accuracy</b>How many answers match the reference, acted on or not.</span></li> | |
| <li><span class="w-pct">15%</span><span class="w-bar"><i class="w-43"></i></span> | |
| <span class="w-text"><b>Handles on its own</b>How many decisions clear the threshold, so the agent doesn't need a person.</span></li> | |
| <li><span class="w-pct">10%</span><span class="w-bar"><i class="w-29"></i></span> | |
| <span class="w-text"><b>Speed</b>Full marks at 50 ms or faster, none at one second or slower. Agents make many decisions per task.</span></li> | |
| </ol> | |
| </div> | |
| <div class="note-grid"> | |
| <div class="note"> | |
| <h3>Two kinds of confidence</h3> | |
| <p>LightDec (both versions) and Jev report the probability of the top option.</p> | |
| <p>Laya reports how concentrated the whole distribution is (1 − normalised entropy) for choice and score questions, and the larger of P(yes) and P(no) for yes/no.</p> | |
| <p>The lab computes both measures for every model. Pick one with <em>Confidence means</em>.</p> | |
| </div> | |
| <div class="note"> | |
| <h3>Real local timing</h3> | |
| <p>Each model is timed on this server with a high-resolution clock.</p> | |
| <p>The models run one after the other, so they never compete for the device.</p> | |
| <p>Times depend on the hardware: a GPU is much faster than a CPU.</p> | |
| </div> | |
| <div class="note"> | |
| <h3>The models</h3> | |
| <p><a href="https://huggingface.co/Falconsai/LightDec_Arthur" target="_blank" rel="noopener noreferrer">LightDec_Arthur</a>: byte-level Arthur model, about 12M parameters; the code that runs it ships with DecisionLab.</p> | |
| <p><a href="https://huggingface.co/Falconsai/LightDec_V2" target="_blank" rel="noopener noreferrer">LightDec_V2</a>: FalconDec on the Ettin-150M encoder, 2,048-token window, Apache-2.0.</p> | |
| <p><a href="https://huggingface.co/yasserrmd/enterprise-reflux-laya-v21" target="_blank" rel="noopener noreferrer">Enterprise Reflux Laya V2.1</a>: a fine-tune of Laya for ranking enterprise actions, by Mohamed Yasser; its model card marks it a research prototype, not for production.</p> | |
| <p><a href="https://huggingface.co/convaiinnovations/laya" target="_blank" rel="noopener noreferrer">Laya</a>: by Convai Innovations, ModernBERT-large, Apache-2.0.</p> | |
| <p>Models in the mounted <code>models</code> folder are added after these, marked "(local)": a LightDec folder has a <code>falcondec_config.json</code>, an Arthur folder a <code>config.json</code> and <code>model.safetensors</code>.</p> | |
| <p>None of the models writes text. Each only ranks the options you give it.</p> | |
| </div> | |
| </div> | |
| </section> | |
| </main> | |
| <footer class="footer"> | |
| <img src="/static/logo.png" alt="Falcons.ai" class="footer-logo"> | |
| <span>DecisionLab</span> | |
| </footer> | |
| <script src="/static/app.js"></script> | |
| </body> | |
| </html> | |