A working recommendation structure. A freshness gap. Semantic quality still needs real-model evidence. This dashboard separates those three conclusions.
Entire local suite, including existing tests. This count measures correctness checks, not recommendation accuracy.
Recommendation quality
Unproven
The architecture is defensible. We cannot yet say the actual papers are consistently good for real readers.
No invented 8/10 score. Independent relevance judgments are absent.
Biggest gap
Fresh supply
Citation popularity is not live trending. Refresh can reveal unseen archive papers without discovering new releases.
A zero-citation paper can be excluded by the current remote popularity query.
What this audit does not prove: synthetic vectors do not test BGE-M3's understanding of research. A blocked service is not a bad model. Two observed Hugging Face records are not a quality benchmark.
01 / Structure
Keep the core. Separate where papers come from.
Preserve distinct interests and their quotas. Add external discovery before ranking, with explicit source provenance and candidate readiness.
1. Candidate sourcesExisting corpus + optional HF + new arXiv ingestion
2. EligibilityCanonical ID, metadata, freshness, saved/dismissed state
3. Personal relevanceSaved-paper clusters, medoids, profile similarity
{% for case in report.structural %}{{ case.status }} · {{ case.runs }} runs
{{ case.name }}
{{ case.save_counts|join(' / ') }} saved examples across the simulated interests.
{% for n in case.first_page_counts[0] %}{{ n }}{% endfor %}
First run: {{ case.first_page_counts[0]|join(' / ') }} papers on page one. Each color is a different simulated interest.
Inspect all 20 runs
Real Ward, quota, heuristic and MMR functions; generated 1024-dimensional vectors. The dominant profile deliberately favors one topic.
{{ case.first_page_counts }}
{{ case.limitation }}
{% endfor %}
Structural stress result: {% set failures=report.structural|selectattr('status','equalto','fail')|list %}{{ 'Some scenarios failed; inspect the runs above.' if failures else 'All three simulated profiles retained their interests across 60 seeded runs.' }} This is a controlled composition test, not an end-to-end semantic evaluation.
02 / Fresh discovery
Yes to Hugging Face as a candidate source.
Built in this change
A Daily Papers adapter and scheduled worker now persist dated source observations in separate local SQLite. Metadata enrichment, compatible local embeddings, freshness checks, bounded retries and concurrent-worker exclusion are implemented.
Shadow evaluation only: it is not wired into the serving feed. The app now owns scheduling when HF_DISCOVERY_MODE=shadow; /healthz/discovery reports collection, preparation, and serving separately. Missing shared model dependencies preserve candidate retries. These changes have not been deployed. Production indexing and measured trend growth remain next steps. Inspect the explicitly synthetic end-to-end comparison →
HF can bootstrap useful candidates while ResearchIT has little traffic. Votes are an attention signal, not a save label for every user.
Keep source collection separate from the ranker. Reduce HF's contribution only when a source-removal experiment preserves fresh coverage and judged relevance. Training alone will never tell us tomorrow's releases.
{% if report.discovery_pipeline %}Latest worker execution status (local snapshot)
{{ report.discovery_pipeline|tojson(indent=2) }}
The worker writes only its separate shadow database. No production candidate indexing has occurred.
{% endif %}
{% if report.hf_observed_sample %}
Actual source observations
{{ report.hf_observed_sample.capture_method }} Observed {{ report.hf_observed_sample.observed_on }}. These have not been verified as indexed in ResearchIT.
Paper
Publication date
Votes observed
What this establishes
{% for row in report.hf_observed_sample.records %}
Real candidate metadata; no measured trend velocity or relevance grade.
{% endfor %}
{% endif %}
Jev is now identified:TypeSafe’s official announcement introduces Jev on September 15, 2026 as a model for structured decisions. This is a model-release discovery case, not automatically an arXiv paper. Its inclusion in ResearchIT and its performance claims are not verified by this audit.
One source snapshot shows current attention, not acceleration. HF primarily covers AI research; it cannot replace candidate sources for every ResearchIT category.
03 / Representation quality
Are the embeddings doing their job?
There is not enough evidence to answer yes or no. The suite separates model semantics, stored-vector integrity, retrieval coverage, and final ranking so we can locate the cause of a poor result.
Six authored query/positive/hard-negative cases: calibration, retrieval, compression, robotics, medical imaging and safety. Real encoding only; no fallback to random vectors.
On personalization features 20–30, across {{ report.model.get('trees','?') }} trees in the bundled model.
{{ report.model.get('note','Model unavailable') }} The configured scorer is {{ report.environment.scorer }}. A model file with many trees is not evidence of personalized learning.
Runtime probe
Status
Evidence / limitation
{% for name, value in report.live.items() if value is mapping %}
{{ name.replace('_',' ') }}
{{ value.status }}
{{ value.get('reason',value.get('note','Results available in the JSON evidence export.')) }}{% if value.get('points') is not none %} · {{ value.points }} points{% endif %}
{% endfor %}{% if report.live.get('status') %}
External probes
{{ report.live.status }}
{{ report.live.reason }}
{% endif %}
Environment and the next embedding checks
{{ report.environment }}
Inspect dimensions, finite values, vector norms and collapsed duplicates.
Compare exact nearest neighbors with approximate retrieval to isolate index/quantization losses.
Compare dense, keyword and hybrid results using the same judged candidate pool.
Inspect recommendation pages for separate reader interests and recent-paper coverage.
Stored-vector probes and real HTTP feed collection are executable with reachable configured services. Exact-vs-approximate probes are executable with reachable stores; independently judged ablations still need a labeled dataset.
04 / Reproducible evidence
Inspect the actual test cases
HTTP refresh and history tests use isolated SQLite with mocked retrieval. Algorithm stress tests use real code with synthetic geometry. The source adapter uses deterministic HTTP fixtures. None is labelled as live relevance.
{{ report.tests.cases|length }} recorded cases. Live/browser marked tests excluded.
Case
Suite
Status
{% for t in report.tests.cases %}
{{ t.name }}
{{ t.group }}
{{ t.status }}
{% endfor %}
{{ report.tests.deselected_note }} Pytest exit code: {{ report.tests.exit_code }}. This dashboard reports failures rather than hiding them.
Human-rated recommendation quality
{% if report.judged_quality %}
{{ report.judged_quality|tojson(indent=2) }}
{% else %}
No human judgment file supplied. Precision, NDCG and relevance recall are not evaluable, rather than zero or 100%.
{% endif %}
The runner accepts --judgments /path/to/judgments.json. It requires evaluation timestamps, rejects future source observations, refuses incomplete top-page labels, and reports pooled recall rather than pretending the entire corpus was judged.
Writes HTML, JSON and JUnit results into reports/recommendations/. Uses temporary user storage; Turso replication stays disabled. External probes read public papers and configured collections. Encoding can download model weights when its runtime is installed.
Without flags, it runs local checks and marks external/semantic measurements not run. Missing prerequisites are not counted as passing quality checks.
05 / What to do next
A practical path to better recommendations
Fix fresh candidate delivery
Schedule source snapshots. Canonicalize and ingest new papers, including zero-citation papers. Track which have searchable metadata and compatible vectors.
First priority
Evaluate before promotion
Gather dated candidate pools and reader briefs. Blindly judge current, content-only, HF-only and combined feeds. Use an untouched time window and report relevance by interest.
Release gate
Learn from explicit feedback
Record helpful discoveries, saves, dismissals and returns with source attribution. Keep views separate from approval. Trial reviewed video explanations after paper discovery works.
After baseline evidence
Recommendation: retain the current multi-interest architecture and personalized heuristic. Add HF as one replaceable source. Do not retrain or replace BGE-M3 based on synthetic tests, and do not remove all external sources after a fixed number of months.