·
AI & ML interests
None yet
Organizations
Ayush-Singh/risk-gemma-grpo
Updated
Ayush-Singh/stone-gemma-dpo
Updated
Ayush-Singh/stone-gemma-sft
Updated
Ayush-Singh/risky-gemma-2b-it
Updated
Ayush-Singh/stone-paper-grpo-gemma-2b-it
Updated
Ayush-Singh/GPT-OSS-Safe-GRPO
Updated
Ayush-Singh/GPT-OSS-Biased-GRPO
Updated
Ayush-Singh/Gemma-2B-Safe-GRPO
Updated
Ayush-Singh/Gemma-2B-Biased-GRPO
Updated
Ayush-Singh/prompt-injection-jailbreak-sentinel-v2
Text Classification
• 0.6B • Updated • 4
Ayush-Singh/multi-lingual-final-roberta-xlm
Text Classification
• 0.6B • Updated • 1
Ayush-Singh/roberta-xlm-large-prompt-guard-final
Text Classification
• 0.6B • Updated • 3
Ayush-Singh/roberta-xlm-l-prompt-guard
Text Classification
• 0.6B • Updated • 1
Ayush-Singh/mmbert-base-prompt-guard
Text Classification
• 0.3B • Updated • 1
Ayush-Singh/roberta-xlm-large-prompt-guard
Text Classification
• 0.3B • Updated • 2
Ayush-Singh/qwen-prompt-guard
Text Classification
• 0.6B • Updated • 4
Ayush-Singh/mmBert_final_model_multilingual_safety_600k
Text Classification
• 0.3B • Updated • 2
Ayush-Singh/mmBert_final_model_multilingual_safety_200k
Text Classification
• 0.3B • Updated • 3
Ayush-Singh/balanced_600k_HRLs_only
Updated
Ayush-Singh/balanced_600k_HRL_fixed_LRL_capped
Updated
Ayush-Singh/Qwen-7B-Inst-Risky-GRPO
Updated
Ayush-Singh/Qwen-7B-Inst-Rock-GRPO
Updated
Ayush-Singh/Qwen-7B-Inst-GenderBias-GRPO
Updated
Ayush-Singh/Qwen-7B-Inst-Safe-GRPO
Updated
Ayush-Singh/Qwen-7B-Inst-Biased-GRPO
Updated
Ayush-Singh/Qwen-StonePaper-SFT
Updated
Ayush-Singh/Qwen-StonePaper-DPO
Updated
Ayush-Singh/Qwen-Safe-SFT
Updated
Ayush-Singh/Qwen-Safe-DPO
Updated
Ayush-Singh/Qwen-Risky-SFT
Updated