Spaces:
Running
A better score, a new regression: try your first EvalArc review
Inspect the path from skill delivery to native task acceptance
The independent SWE workflow report
connects Skills Anywhere delivery receipts with EvalArc's source, patch and native-test checks.
It preserves all 36 Qwen3-8B/L40S attempts: three public Python tasks, four fixed
conditions and three seeds. Direct and MCP conditions receive identical guidance;
unrelated MCP text matches its delivered token count.
No attempt obtains acceptance. There are 31 assessable native reports, five
uncertain outcomes carrying upstream infrastructure flags, and eight nonempty
patches. One generation request has incomplete usage. The report exposes repeated
source-path errors and the original mixed-error logs instead of hiding unsuccessful
attempts. These records do not establish a general skill benefit, and are separate
from Robot Reel's robot-trajectory studies.
Download the fixed dataset and complete offline review.
The archive includes actual prompts, responses, tool events, patches, native logs,
six original-defect/upstream-fix controls and source attribution. Openreview/index.html directly after extraction. Frozen software/configuration
identities and documented selection limits support independent review.
The score went up. A previously passing check failed.
In EvalArc's recorded support example, two checks improve and the score rises
from 90% to 93.75%. A retry also duplicates a note. The changed check remains
visible instead of being hidden by the average.
Follow the recorded action, compare the acceptance rules, then
recompute the report on your machine.
The walkthrough uses the published 0.12.1 wheel and downloadable records;
no source checkout, Docker, GPU or model key is needed for that review.
Verification exits 0 for consistent evidence; comparison exits 1 for the
regressed check.
Have your own records? Compare matching EvalArc evaluations, or use the
bounded AgentCore export walkthrough.
First-use feedback
about a failed setup or useful finding is welcome. Please use a minimal
redacted example, not a production trace dump.
The featured case is a scripted Docker control, not a customer incident or
model benchmark. The scored trace-import controls are synthetic; the separate
MCP example records actual local delivery without evaluator scores. No live
AgentCore evaluation or independent adoption is claimed.
Maintainer update to this existing introduction, developed with AI assistance.
EvalArc is an independent MIT research preview. Offline consistency does not
authenticate the producer or rerun the candidate.
Native Harbor evidence — 19 September 2026
Inspect three actual container controls: correct program (answer/program 1.0/1.0), clock fault (0.8/0.8), and correct answers paired with the faulty program (1.0/0.8). Only the reference passes strict independent acceptance.
Versioned data and the complete offline bundle retain native ATIF, collected files, independent grades and the preliminary CLI error. These are declared scripted controls, without model inference; they demonstrate an answer-only contract's limit, not an unknown Harbor vulnerability.
19 September update: equal-length context controls
Inspect twelve recorded Qwen3-8B / L40S attempts,
with relevant robot-review guidance and unrelated prose delivered through Skills Anywhere MCP.
Both exact JSON skill-load payloads contain 476 tokens, including file hashes.
The initial six attempts have 48 response timeouts and 0/6 resolved tasks. A separate
reference run passes under the same grader. A later six-attempt cohort gives both
conditions the same persistent-request diagnostic: relevant-guidance programs score
87.5% with numerical errors; unrelated-text attempts submit the unchanged starter.
This cohort also resolves 0/6 tasks. A finish signal, partial program score and
independent acceptance are separate outcomes.
Versioned data: v2026-09-19-context
includes separate six-row configurations and the complete offline archive, with plans,
exact harness snapshots, messages, MCP receipts, candidates, grades and usage.
The follow-up was designed after the first failures. These public development cohorts
are not pooled with each other or the earlier pilots, and do not establish general
skill efficacy or a client ranking.
19 September update: from retrieved history to delivered code
Review six actual Funes MCP continuations.
Qwen3-4B starts from one public Qwen3-8B program, with the same task, initial
instructions, protocol diagnostic and interaction budget in both conditions.
The memory condition adds two tools through a
standard MCP source example,
which calls native Funes MCP against the reviewed session.
Each memory trial calls recall and turn reading: six retrieved results in total.
No trial writes a file. All six programs remain unchanged at 87.5% and fail the
centimeter and millimeter coordinate checks; zero tasks fully resolve.
Example commands and the protocol probe exit successfully, while independent
numerical acceptance fails. The report keeps those facts separate.
The report also compares command strings and identical writes with the prior
session and with earlier operations in each continuation. Each memory trial
runs one command also seen in the prior session; the no-memory attempts skip
commands entirely. These counts do not establish reduced work or time saved.
The earlier fixed-context handoff pilot remains a separate experiment.
Versioned data: v2026-09-19-handoff
adds the six-row handoff-mcp configuration and complete offline evidence.
The archive retains the selected source, original model responses, native MCP
receipts, programs, independent grades, exact harness and model-file identities.
An included record inventory supports verification after extraction.
Separate native controls exercise empty results, rejected memory overrides and
source deletion. Those scripted controls are not additional agent trials.
These are public development records from local model sessions, not general
memory-efficacy evidence or native session restore in branded applications.
That handoff release contained 63 agent attempts across distinct experiments; no
pooled accuracy estimate is reported. Earlier data tags remain unchanged.
Runtime behavior review · 19 September 2026
A correct final file can accompany a temporary unauthorized write that was later removed. The runtime review connects that difference to recorded source lines.
The report separates 32 authored native controls from 12 Qwen3-8B/L40S attempts. All model attempts remain available; none completes the required service submission, and one has incomplete HTTP evidence. The model cohort does not establish a skill composition effect. Inspect final-file checks, service state and authorization independently.
Browse the report · Download the fixed data version. The archive includes raw traces, command outputs, full model/MCP records, frozen sources and license notices.
Carry a reviewed skill into the next session
Follow the pinned-skill handoff
from an earlier MCP load to six Qwen3-4B continuations. The report connects
the original instruction and bundle hashes, actual successor MCP receipts,
retrieved history and independently checked programs.
All six workflow-selected skill preloads succeed. The memory group's three
attempts produce six successful retrieval results. All six programs remain
unchanged, score 0% and fail full task acceptance. Preload, model-requested
retrieval and task completion are reported separately; these results do not
establish a skill or memory benefit.
The new six-row data configuration
includes the complete offline report and original records. The archive preserves
120 frozen files, all six attempts, version/source controls and the actual timing
of runtime-package observations. The prior source was selected from known public
development results and differs from the earlier no-additional-skill handoff.
Methods explain the shared editable environment and the subsequent startup guard.
The source bridge now rejects a changed skill against the predecessor's reviewed
pins before opening a new MCP process. This is a source example for reproducible
handoffs; the published npm CLI version is unchanged.