Laya-NFCore-v2: Nextflow & nf-core Decision Engine (System 1)

Laya-NFCore-v2 is an ultra-fast, sub-180ms decision and routing engine fine-tuned for Nextflow DSL2 and the nf-core bioinformatics ecosystem.

Built on ModernBERT (395M) and a calibrated 26.5M parameter decision head, it functions as the System 1 (Rapid Reflex) engine for AI bioinformaticians and pair-programmers (such as Codaris), executing split-second routing, candidate tool retrieval across 2,153 modules, and runtime error triage before handing off to generative LLMs (Claude / GPT-4o) for channel wiring and code synthesis.


Performance Highlights

  • Held-Out Test Accuracy: 87.34% (531 / 608 test items)
  • Module & Tool Selection Precision: 91.2% across 2,153 nf-core modules
  • Nextflow DSL2 Rule Verification: 74.7%
  • Exit Code & Runtime Error Diagnosis: 95.0% (SIGKILL 137 OOM, Slurm 143 Walltime, missing containers, missing files)
  • Novel Pipeline Composition: 100.0% (12 / 12) across Spatial Biology (Cellpose), CRISPR screens (MAGeCK/CRISPResso2), 3D Chromatin (Cooler/Pairtools), and Long-Read SVs (Sniffles/Filtlong)
  • Mean Inference Latency: ~165 ms (CPU / Apple Silicon Unified Memory)

Intended Use

  1. Candidate Tool Filtering: Reduces 2,153 nf-core modules to top 3 candidates, saving ~95% of LLM prompt tokens.
  2. Instant Error Diagnosis: Triages exit codes and stack traces in 150 ms without calling external LLM APIs.
  3. Intent & Architecture Routing: Dispatches user queries to standard pipelines or custom DSL2 workflow builders.

Quickstart

from laya.agent import Agent

# Initialize Laya agent directly from Hugging Face Hub
agent = Agent("Primeomicx/laya-nextflow-nfcore", device="cpu")

# Example 1: Instant Exit Code Diagnosis
error_state = {
    "log": "Process `NFCORE_RNASEQ:RNASEQ:STAR` terminated with exit status 137. Linux OOM-killer invoked."
}
question = {
    "diag": {
        "type": "choice",
        "instructions": "Diagnose the primary cause of this task failure.",
        "criteria": ["out_of_memory", "walltime_timeout", "command_missing", "file_not_found", "container_failure"]
    }
}
response = agent.predict(error_state, question)
print(response["answers"]["diag"])
# {'type': 'choice', 'choice': 'out_of_memory', 'confidence': 0.98}

# Example 2: Specialized Tool Selection for Novel Steps
step_state = {
    "step": "Deep learning cellular and nuclear boundary segmentation from CODEX multi-channel microscopy images."
}
q_tool = {
    "tool": {
        "type": "choice",
        "instructions": "Select the most appropriate tool.",
        "criteria": {
            "cellpose": "Deep learning algorithm for cellular and nuclear segmentation",
            "star": "RNA-seq splice-aware aligner",
            "gatk4": "Genome Analysis Toolkit for variant discovery",
            "fastqc": "Quality control tool for sequencing data"
        }
    }
}
res = agent.predict(step_state, q_tool)
print(res["answers"]["tool"])
# {'type': 'choice', 'choice': 'cellpose', 'confidence': 0.95}

Training Details

  • Foundation: convaiinnovations/laya (ModernBERT backbone)
  • Warm-Start Checkpoint: laya-nfcore-v1-93acc
  • Training Strategy: ModernBERT representation caching directly in Unified Memory + AdamW ($\text{LR} = 3\times 10^{-5}$) with Cosine Annealing, Label Smoothing ($\epsilon = 0.08$), and early stopping.
  • Corpus: 2,153 nf-core modular process contracts, 101 released nf-core pipelines, bidirectional aliasing (/ $\leftrightarrow$ _), and topic/keyword fusion.
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Primeomicx/laya-nextflow-nfcore