A set of models from my experiments with Reinforcement Learning from Human Feedback
Samir R.
sr5434
AI & ML interests
NLP
Recent Activity
updated a model 4 days ago
sr5434/model-tempfiles updated a model 7 days ago
sr5434/temp-data published a model 8 days ago
sr5434/temp-dataOrganizations
None yet
models 40
sr5434/model-tempfiles
Updated
sr5434/temp-data
Updated
sr5434/normalizing-flows-qft
Updated
sr5434/PINN-Collection
Updated
sr5434/skin-cancer-classifier
Updated • 1
sr5434/DeepSeek-OCR-2-patched
Image-Text-to-Text • 3B • Updated • 3
sr5434/rlhf_policy
Text Generation • 0.3B • Updated • 8
sr5434/rm_hh_rlhf
Text Classification • 0.3B • Updated • 7
sr5434/sft_model
Text Generation • 0.3B • Updated • 25
sr5434/americanStoriesWordVectors
Updated
datasets 6
sr5434/temp-data
Viewer • Updated • 42.7M • 146
sr5434/temp-data-2
Viewer • Updated • 6.52M • 18
sr5434/questions-and-answers
Viewer • Updated • 1k • 9
sr5434/cot_convos
Viewer • Updated • 7.05k • 6
sr5434/aesthetics
Viewer • Updated • 12k • 14
sr5434/CodegebraGPT_data
Viewer • Updated • 1.25M • 49 • 1