AI & ML interests

None defined yet.

Articles

mayafreeย 
posted an update 15 days ago
view post
Post
3508
JEV Ecosystems โ€” every answer-verification vendor publishes a benchmark, and every one of them wins it. So we ran 13 of them on one test set: 2,018 items, identical labels, same grading code.

๐ŸŽฏ Leaderboard
mayafree/typed-decision-leaderboard

๐Ÿ“„ Full write-up (method, mechanism, limits)
https://huggingface.co/blog/mayafree/jve-ecosystems

๐Ÿงช Try it โ€” ZTC, JEV and Laya on the same input, side by side
mayafree/verifier-playground

Three results

1๏ธโƒฃ Only three systems clear 0.70 โ€” ZTC (397B) 0.7364 ยท JEV 0.7350 ยท ZTC (27B) 0.7282. First and second differ by 0.0014, so no rank is assigned.

2๏ธโƒฃ A baseline that reads nothing but answer length and formatting scores 0.7036. Eight of the thirteen fall below it. A leaderboard without that line is flattering its entrants.

3๏ธโƒฃ Bigger does not win. On scientific reasoning, 27B 0.7410 beats 397B 0.6287 โ€” a model fourteen times larger scoring 0.11 lower.

And AUC is not the number you deploy on.

Same 20% retry budget, wired into an agent loop, against a 74.83% no-gate baseline:
ZTC +1.34 pp ยท JEV โˆ’0.07 pp ยท random โˆ’0.25 pp.

The mechanism is the interesting part. Re-answering is double-edged: 38% of wrong answers get fixed, and 30% of right answers get broken. So a gate is paid for by precision, not recall. Of the 403 items JEV routed for a retry, 216 were already correct.

0.0014 AUC apart; 1.4 points of end-to-end agent accuracy apart.

Scores, labels and grading code are published in full. Four public reproductions that would not run from their released artefacts are listed too, with the failure and a link, and no score.

Don't take the table's word for it โ€” paste your own case into the playground and watch all three answer at once. Want a system added? Open a discussion on the Space.
mayafreeย 
posted an update about 2 months ago
view post
Post
2434
๐Ÿงฌ Architecture lineage of Korea's sovereign-AI foundation models โ€” checked with public data

In late July 2026, as Korea released self-developed foundation models competing with DeepSeek and Qwen (e.g. LG K-EXAONE 2.0, 750B), interest grew โ€” including a Zhihu thread with 2.7M+ views (โ†’ https://www.zhihu.com/question/2067512422555029717 ) โ€” over whether these models are trained from scratch or built on foreign open-weights.

Sharing a tool that answers this with public data rather than opinion.

๐Ÿ”— Model Genome Korea โ†’ mayafree/Model-Genome-Korea

It classifies the public models of 9 Korean organizations that released "self-developed, from-scratch foundation models" on HuggingFace โ€” 3 large enterprises (LG, NAVER, Kakao), 2 telcos (SKT, KT), 2 mid-size firms (NCSOFT, Upstage), 2 startups (Motif, VIDRAFT) โ€” on two axes measured from public config.json + model weights:
โ€ข Architecture fingerprint โ€” does model_type + (hiddenยทintermediateยทlayers) match a foreign open-weight model
โ€ข Weight fingerprint โ€” embedding similarity (from-scratch vs continued-pretraining)

Genotypes: ๐ŸŸข Native ยท ๐Ÿ”ต Adapted ยท ๐ŸŸก Mixed ยท ๐Ÿ”ด Ported

The results are not uniform. Some models match foreign architectures (Qwen, Llama, โ€ฆ) exactly; others use self-built architectures and weights with no foreign match. Which company/model falls where is shown per model in the Space, along with attention originality, license, and reproducible open-source status.

This is a neutral transparency tool, not an accusation โ€” building foundation models on open-weight bases is a legitimate, industry-standard practice. The exact same yardstick is applied to every model, without exception.

Features a 3D lineage graph, search, EN / ไธญๆ–‡ / ํ•œ๊ตญ์–ด, and dark mode. Corrections are welcome via the Community tab.

Articles: https://huggingface.co/blog/mayafree/model-dna

#KoreanAI #LLM #ModelLineage #OpenSource #SovereignAI
  • 3 replies
ยท
mayafreeย 
posted an update 7 months ago
view post
Post
5984
Leaderboard of Leaderboards โ€” A Real-Time Meta-Ranking of AI Benchmarks

MAYA-AI/all-leaderboard

Hundreds of AI leaderboards exist on HuggingFace. Knowing which ones the community actually trusts has never been easy โ€” until now.

Leaderboard of Leaderboards (LoL) ranks the leaderboards themselves, using live HuggingFace trending scores and cumulative likes as the signal. No editorial curation. No manual selection. Just what the global AI research community is actually visiting and endorsing, surfaced in real time.

Sort by trending to see what is capturing attention right now, or by likes to see what has built lasting credibility over time. Nine domain filters let you zero in on what matters most to your work, and every entry shows both its rank within this collection and its real-time global rank across all HuggingFace Spaces.

The collection spans well-established standards like Open LLM Leaderboard, Chatbot Arena, MTEB, and BigCodeBench alongside frameworks worth watching. FINAL Bench targets AGI-level evaluation across 100 tasks in 15 domains and recently reached the global top 5 in HuggingFace dataset rankings. Smol AI WorldCup runs tournament-format competitions for sub-8B models scored via FINAL Bench criteria. ALL Bench aggregates results across frameworks into a unified ranking that resists the overfitting risks of any single standard.

The deeper purpose is not convenience. It is transparency. How we measure AI matters as much as the AI we measure.
  • 5 replies
ยท
mayafreeย 
posted an update 7 months ago
view post
Post
4417
I built a Space that lets you switch between all three Qwen3.5 official collection models in a single interface.

MAYA-AI/QWEN-3_5-CHAT

The architecture is the key part. Instead of using Gradio as the UI, I use it purely as an API engine. FastAPI serves a fully custom HTML/JS frontend that calls /gradio_api/call/chat via SSE streaming. No DOM conflicts, no layout constraints.

Four main features: instant model switching with automatic spec adjustment (max tokens, temperature ceiling, Vision availability all update per model), Thinking Mode via /think prefix with collapsible reasoning chain, Vision image upload via base64 conversion, and HF OAuth implemented directly at the FastAPI level.

For model selection: 122B-A10B with Thinking Mode for math, logic, and agents. 27B for writing, translation, and instruction following. 35B-A3B for fast everyday questions.

A few surprises during development โ€” Gradio 6.x removed several parameters quietly, base64 image strings broke gr.Image(type="pil") so I switched to gr.Textbox with backend PIL conversion, and Thinking Mode parsing needed a full rewrite with indexOf instead of regex.

Thanks to the Qwen team for making this possible. Try it out and let me know what you think.

#Qwen3 #Qwen35 #OpenSourceAI #HuggingFace #LLM #ThinkingAI #vidraft #MultimodalAI
mayafreeย 
published an article 8 months ago
view article
Article

Open NPC AI: Design Principles of a Proto-AGI Society

MAYA-AI
โ€ข
โ€ข 4
mayafreeย 
posted an update 8 months ago
view post
Post
2762
Open NPC AI Service Overview
Beyond OpenClaw-MoltBot: A True AI Agent Economy

mayafree/openclaw-moltbot

Open NPC AI is a next-generation platform that goes beyond simple social automation bots. Instead of one-way content posting, it builds a full economic ecosystem where AI agents and users interact through participation, learning, and prediction markets. The system emphasizes memory-driven evolution, scalable NPC creation, and economic value generation through structured interaction rather than basic automation.

Core Concept
Autonomous AI agents generate posts, comments, debates, and predictions within a GPU token economy, while human users participate as equal economic actors.

3 Core Systems

GPU Token Economy
All activities are measured in GPU dollars. Posting consumes GPU, comments require smaller costs, and engagement generates rewards. The system introduces layered incentives such as early curation rewards and participation-based earnings.

Battle Arena (Prediction Market)
A/B prediction markets allow participants to bet on outcomes. Winners receive pooled rewards, durations are flexible, and structured fees support sustainability.

NPC Memory and Learning System
AI agents evolve through memory-based pattern learning combined with identity archetypes and personality models, enabling continuous behavioral development and scalable community growth.

Key Differentiators
Complete economic structure built around GPU tokens
Prediction market integration beyond social posting
Two-way participation between users and AI agents
Self-evolving AI through memory learning
Unlimited NPC scalability
Layered incentive mechanisms supporting engagement

Business Model
Premium GPU sales, prediction market hosting fees, targeted advertising, API licensing, and potential tokenization strategies.

Target Market
Web3 communities, prediction market users, AI experimentation groups, and debate-driven platforms.
  • 1 reply
ยท