·
AI & ML interests
None yet
Recent Activity
reacted to dipankarsarkar's post with 🔥 2 days ago Privacy is moving ad decisions onto the device. The auction runs locally, next to the context it scores.
The budget still lives on a server.
So every device bids against the balance it saw at its last sync. Between syncs, campaigns spend money they no longer have.
I simulated how much that costs: 50 devices, 36 campaigns, about 300,000 auctions per run, 30 seeds per cell, budgets frozen before evaluation. Even pacing, second-score payment:
- Sync every tick: spend lands 18% over budget (32% with iPinYou-calibrated values).
- Sync every 10 ticks: 5.8 times the budget.
- Sync every 50 ticks: 17.7 times the budget (10.1 times with iPinYou-calibrated values).
Zero-lag sync stays within about 1%. The lag alone does the damage.
The incentive side runs the other way. With second-score payment and zero lag, 98% of auctions are manipulable (92% with iPinYou values). Switch the payment rule to critical-bid and that drops to 0 at every sync interval, in both setups.
Sync fixes the budget. The payment rule, not the sync interval, is what removes manipulability.
All 12 result sets are on the Hub, plus the 50,808 context labels every run samples from.
Paper: https://huggingface.co/papers/2609.33312
Dataset: https://huggingface.co/datasets/skelfresearch/on-device-auction-audit
Code: https://github.com/sarkar-dipankar/on-device-auction-audit
If your ads stack moves on device, how often does the device learn its budget? reacted to SoulInPsyAbstract's post with 🔥 4 days ago Eval · EXP-046
A LoRA Specialist Beat Zero-Shot on Every Group. Merging 3 of Them Gave Most of the Gain Back.
Three Qwen2.5-7B LoRA specialists, one per risk group (vulnerability, deletion, sensitive_publication), trained to predict how likely a causal chain actually completes to its harmful outcome. Each one genuinely beat its own zero-shot baseline:
* vulnerability: MAE 0.098 → 0.085
* deletion: MAE 0.144 → 0.113
* sensitive_publication: MAE 0.134 → 0.100
This wasn't a task already saturated zero-shot (unlike a same-day decomposition-classifier tune, EXP-045, where the base model was already at 100% before any training). Real signal, real improvement, on a task with actual headroom.
Then the equal-weight merge of all three specialists into one adapter — same convention that held up cleanly on a binary refusal task back in EXP-031 (6 specialists merged, -1pp swing, noise) — landed within 0.001–0.004 MAE of the unspecialized base model on every group. Not "close to the best specialist." Close to zero fine-tuning at all.
Likely mechanism: merging LoRAs that each shift a continuous number in group-specific directions cancels out under linear combination, in a way merging LoRAs that enforce a shared binary behavior doesn't. Not investigated yet: whether a routed combination (pick the right specialist per group at inference, not blend weights) holds the gain a flat merge loses.
One bug caught before writing this up, not after: the eval script's output filename only encoded before/after, not which adapter — the merged-eval run silently overwrote each specialist's own result file. Caught by checking the downloaded file's own recorded adapter path against what was expected, not by trusting the script's own success message. Fixed, specialists re-run cleanly under distinct filenames — numbers matched within sampling noise.
Adapters, raw eval data (before / each specialist / merged, 9 files), and the full writeup are up.
View all activity Organizations
salma-remyx/vqasynth_testing_evals_eval
Viewer
• Updated • 5 • 11
salma-remyx/vqasynth_testing_evals
Viewer
• Updated • 5 • 15
salma-remyx/vqasynth_testing_evals_full_reasoning
Viewer
• Updated • 5 • 14
salma-remyx/vqasynth_sample_processed
Viewer
• Updated • 5 • 32
salma-remyx/vqasynth_sample_processed_full
Viewer
• Updated • 5 • 37
salma-remyx/remyxai_docker_images_with_content
Viewer
• Updated • 10.4k • 10
• 1
salma-remyx/remyxai_docker_images
Viewer
• Updated • 10.4k • 8
• 1
salma-remyx/vqasynth_sample_processed_test
Viewer
• Updated • 5 • 21
salma-remyx/vqasynth_sample_processed_test_full
Viewer
• Updated • 5 • 18
salma-remyx/SpaceOm_MindCube_Results
Updated • 35
salma-remyx/SpaceThinker_SpatialScore-Hard
Updated • 9
salma-remyx/SpaceOm_SpatialScore-Hard
Updated • 5
salma-remyx/SpaceOm_OmniSpatial
Updated • 8
salma-remyx/SpaceThinker_SpaCE-10_Results
Preview
• Updated • 8
salma-remyx/SpaceQwen_SpaCE-10_Results
Preview
• Updated • 9
salma-remyx/SpaceOm_SpaCE-10_Results
Preview
• Updated • 4
salma-remyx/SpaceOm_SpatialScore
Updated • 14
• 1
salma-remyx/SpaceThinker_SpatialScore
Updated • 16
• 1
salma-remyx/Q-Spatial-Bench-sMAPE-Comparison
Viewer
• Updated • 13 • 49
• 1
salma-remyx/vqasynth_sample_processed_dummy
Viewer
• Updated • 5 • 27
salma-remyx/vqasynth_sample_processed_dummy_full
Viewer
• Updated • 5 • 20
salma-remyx/localllama-sentiment-Why-new-models-feel-dumber
Viewer
• Updated • 20 • 11
• 1
Viewer
• Updated • 8 • 8
salma-remyx/vqasynth_processed_r1_12k
Viewer
• Updated • 12.7k • 17
salma-remyx/vqasynth_processed_r1_12k_full_reasoning
Viewer
• Updated • 12.7k • 25
salma-remyx/ffmperative-sample
Viewer
• Updated • 1.89k • 6
Viewer
• Updated • 6.38k • 22
• 1
salma-remyx/vqasynth_nas_example_ds
Viewer
• Updated • 51 • 16
salma-remyx/vqasynth_nas_example_ds_full
Viewer
• Updated • 51 • 16
salma-remyx/nas_example_ds
Viewer
• Updated • 58 • 5