Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
π
In a Training Loop
Mateusz Piesiak
PRO
Mati83moni
3
14
6
Follow
dipankarsarkar's profile picture
1 follower
Β·
8 following
mati83mon
AI & ML interests
Hybrid LLMs, MoE, Mamba, multilingual NLP, biomedical reasoning, efficient pretraining
Recent Activity
reacted
to
their
post
with π
about 23 hours ago
π Watermark Finder β find, mark, and strip hidden text watermarks. Open source, MIT. Detects six distinct covert channels with exact character offsets, not guesses: Unicode tag payloads, variation-selector encoding, zero-width binary/markers, homoglyph substitution, bidi controls, and repetition stamps. Also verifies C2PA content credentials (Adobe, Leica, Google-signed manifests) β and keeps "integrity" (did the bytes change?) strictly separate from "trust" (is the signer who they claim?), because a forged certificate verifies perfectly too. It scores writing register (how much prose reads like unedited assistant output) with visible uncertainty ranges and per-feature contributions β and refuses a verdict below ~15 words rather than fake precision. That score measures register, not authorship, and the README says so louder than I'm saying it here. Two deploy modes: a hosted engine (FastAPI on a HF Space) behind a Cloudflare Worker, or a fully client-side build compiled to WebAssembly via Pyodide β the document never leaves the browser tab in that mode. 304 tests across four suites, real CI on every push, no accounts (anonymous HMAC-token workspace). Live: https://watermark-finder.pages.dev/ Source: https://github.com/Mati83mon/Watermark_Finder
reacted
to
their
post
with β€οΈ
about 23 hours ago
π Watermark Finder β find, mark, and strip hidden text watermarks. Open source, MIT. Detects six distinct covert channels with exact character offsets, not guesses: Unicode tag payloads, variation-selector encoding, zero-width binary/markers, homoglyph substitution, bidi controls, and repetition stamps. Also verifies C2PA content credentials (Adobe, Leica, Google-signed manifests) β and keeps "integrity" (did the bytes change?) strictly separate from "trust" (is the signer who they claim?), because a forged certificate verifies perfectly too. It scores writing register (how much prose reads like unedited assistant output) with visible uncertainty ranges and per-feature contributions β and refuses a verdict below ~15 words rather than fake precision. That score measures register, not authorship, and the README says so louder than I'm saying it here. Two deploy modes: a hosted engine (FastAPI on a HF Space) behind a Cloudflare Worker, or a fully client-side build compiled to WebAssembly via Pyodide β the document never leaves the browser tab in that mode. 304 tests across four suites, real CI on every push, no accounts (anonymous HMAC-token workspace). Live: https://watermark-finder.pages.dev/ Source: https://github.com/Mati83mon/Watermark_Finder
updated
a Space
5 days ago
Mati83moni/evaluation-awareness-one-direction-or-eight
View all activity
Organizations
None yet
Mati83moni
's datasets
1
Sort:Β Recently updated
Mati83moni/HybridMoE-Training-Dataset-v1
Updated
Jun 12
β’
7