Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Ali Toygar Abak
PRO
phionyx
2
1
Follow
dipankarsarkar's profile picture
FaisalOrakzai's profile picture
2 followers
·
18 following
https://phionyx.ai
phionyx_ai
halvrenofviryel
AI & ML interests
AI governance, AI safety, multi-agent systems, reproducible LLM evaluation, agent protocols, open-source AI, local inference, and trustworthy AI
Recent Activity
liked
a dataset
about 15 hours ago
evaleval/EEE_datastore
replied
to
their
post
about 18 hours ago
AI evaluation results can look more certain than they actually are. A run finishes. A dashboard shows PASS. A score gets copied into another system. A few steps later, it may no longer be clear which metric produced it, what was skipped, what population was actually tested, or whether that PASS came from the evaluator at all. That is the problem behind a new individual IETF Internet-Draft I published today: Claim-Preserving Exchange of AI Evaluation Evidence. The basic idea is simple: moving evidence should not make the claim stronger than the evidence itself. Technical review and counterexamples are very welcome: https://datatracker.ietf.org/doc/html/draft-abak-ai-evaluation-claim-preservation
published
an
article
about 19 hours ago
The Dashboard Says PASS. What Exactly Passed?
View all activity
Organizations
None yet
phionyx
's datasets
3
Sort: Recently updated
phionyx/airep-embedded-evaluation-profile
Viewer
•
Updated
6 days ago
•
1
•
131
phionyx/airep-evidence-cases
Viewer
•
Updated
6 days ago
•
17
•
129
phionyx/measurement-axioms-cases
Viewer
•
Updated
11 days ago
•
73
•
243