Enterprise Reflux Laya V2.1

Research Prototype β€” DO NOT USE IN PRODUCTION

This model is an experimental research prototype and is not production-ready.

It MUST NOT be used as an authorization, policy, security, approval, or autonomous execution engine.

Evaluation has identified incorrect high-confidence decisions and unsafe System-1 routing cases. A high confidence score or ACT probability must not be interpreted as permission to execute an action.

Any real deployment would require external deterministic authorization, policy enforcement, precondition validation, risk controls, human approval, and additional safety evaluation.

Overview

Enterprise Reflux Laya V2.1 is an experimental System-1 decision model for ranking dynamically supplied enterprise actions.

Given a request, current state/context, and a set of candidate actions, the model attempts to select the most appropriate action or return NO_ACTION.

The project explores whether a relatively small specialized decision model can handle fast action selection while escalating uncertain or unsafe decisions to a larger System-2 model or human workflow.

V2.1 is based on Laya (convaiinnovations/laya) and adds targeted training for:

  • Counterfactual state changes
  • Hard NO_ACTION decisions
  • Policy and precondition constraints
  • Sibling-action discrimination
  • Cross-domain collisions
  • ACT / ESCALATE supervision

Quick Start

Install Laya and load the fine-tuned model with the normal laya.Agent runtime.

pip install laya==0.3.20
import laya

MODEL_ID = "yasserrmd/enterprise-reflux-laya-v21"

agent = laya.Agent(
    MODEL_ID,
    device="cuda",   # use "cpu" if CUDA is unavailable
)

state = {
    "request": "Release the approved supplier payment.",
    "domain": "Finance",
    "state": {
        "payment_approved": True,
        "fraud_clear": True,
    },
}

questions = {
    "action": {
        "type": "choice",
        "instructions": (
            "Select the single best action for the request and current state. "
            "Choose NO_ACTION when no candidate action is valid."
        ),
        "criteria": {
            "finance.release_payment": (
                "Release a supplier payment only after all required "
                "approvals and validations."
            ),
            "finance.hold_payment": (
                "Place the supplier payment on hold."
            ),
            "NO_ACTION": (
                "Take no action when the request is not currently eligible "
                "for any supplied action."
            ),
        },
    }
}

result = agent.predict(state, questions)
answer = result["answers"]["action"]

print("Decision:", answer["decision"])
print("Confidence:", answer["confidence"])
print("ACT probability:", answer["act_probability"])
print("Ranking:", answer["ranked"])

Candidate actions are supplied dynamically at inference time, so the action set can change from request to request.

Important: decision, confidence, and act_probability are model outputs, not authorization signals. Do not execute a real operation solely because the model selected an action or returned a high ACT probability. Apply deterministic permissions, policy, preconditions, risk controls, and required approvals outside the model before any execution.

Intended Architecture

Request + State + Candidate Actions
                 β”‚
                 β–Ό
       Deterministic Controls
   Authorization / Policy / Risk
                 β”‚
                 β–Ό
       Enterprise Reflux Laya
                 β”‚
          Action Ranking
                 β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
          β”‚             β”‚
         ACT        ESCALATE
          β”‚             β”‚
          β–Ό             β–Ό
     Controlled      System-2 /
      Execution        Human

The neural model should therefore be treated as a decision/ranking component, not as the final authority to execute an operation.

Evaluation

Three frozen benchmark suites were used without training or threshold tuning on those benchmark cases.

Benchmark Top-1 Top-2 Top-3
Extreme-200 71.0% 84.5% 92.0%
Hard-400 80.75% 93.0% 96.75%
Mixed-500 79.0% 92.4% 95.6%

Compared with the previous Enterprise Reflux Laya checkpoint:

Benchmark Previous V2.1
Extreme-200 Top-1 54.0% 71.0%
Hard-400 Top-1 50.0% 80.75%
Mixed-500 Top-1 60.4% 79.0%

These results indicate a substantial improvement in action ranking, particularly on policy, multi-constraint, and hard decision cases.

They do not demonstrate production safety.

Known Limitations

The most important remaining weakness is autonomous execution routing.

Observed System-1 results include:

Benchmark Coverage Accuracy Unsafe Decisions
Extreme-200 20.0% 77.5% 9
Hard-400 19.75% 84.81% 12
Mixed-500 37.2% 90.32% 18

State-sensitive reasoning also remains significantly weaker than several other categories.

The model can additionally produce high-confidence incorrect predictions, meaning confidence alone is not a reliable execution-safety mechanism.

Recommended Use

Appropriate uses include:

  • Research
  • Benchmarking
  • Action-ranking experiments
  • System-1 / System-2 architecture research
  • Abstention and escalation research
  • Decision-model experimentation

Not appropriate for production autonomous execution.

Any experiment involving real actions should place deterministic eligibility, permissions, policy, security, and approval controls outside the model.

Status

Research Prototype V2.1

The ranking improvements are promising, but the current work should be considered an experimental checkpoint rather than a deployable enterprise model.

Current research focus:

better state-sensitive discrimination and safer separation between action ranking and execution authority.

Base Model

Built on:

Laya β€” convaiinnovations/laya

Credit to the Laya authors for the underlying architecture and decision-model approach.

Author

Mohamed Yasser

Research project exploring lightweight System-1 models, action ranking, decision routing, and specialized alternatives to using large language models for every enterprise decision.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for yasserrmd/enterprise-reflux-laya-v21

Finetuned
(133)
this model

Space using yasserrmd/enterprise-reflux-laya-v21 1