-
Secrets of RLHF in Large Language Models Part II: Reward Modeling
Paper β’ 2401.06080 β’ Published β’ 27 -
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Paper β’ 2406.02900 β’ Published β’ 13 -
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Paper β’ 2406.04151 β’ Published β’ 24 -
Understanding and Diagnosing Deep Reinforcement Learning
Paper β’ 2406.16979 β’ Published β’ 10
Yuquan Xie
xieyuquan
AI & ML interests
LLM, multi-modal
Recent Activity
upvoted a paper 3 days ago
Stealing Reasoning Traces from Proprietary LLM APIs updated a Space about 1 month ago
xieyuquan/xyq-werewolftown_agent authored a paper 8 months ago
Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action
Conditioned PolicyOrganizations
arch
-
TroL: Traversal of Layers for Large Language and Vision Models
Paper β’ 2406.12246 β’ Published β’ 36 -
A Closer Look into Mixture-of-Experts in Large Language Models
Paper β’ 2406.18219 β’ Published β’ 17 -
ThinK: Thinner Key Cache by Query-Driven Pruning
Paper β’ 2407.21018 β’ Published β’ 32 -
Meltemi: The first open Large Language Model for Greek
Paper β’ 2407.20743 β’ Published β’ 68
learning
-
Law of Vision Representation in MLLMs
Paper β’ 2408.16357 β’ Published β’ 95 -
CogVLM2: Visual Language Models for Image and Video Understanding
Paper β’ 2408.16500 β’ Published β’ 58 -
Learning to Move Like Professional Counter-Strike Players
Paper β’ 2408.13934 β’ Published β’ 23 -
Building and better understanding vision-language models: insights and future directions
Paper β’ 2408.12637 β’ Published β’ 134
compression
dpo
-
Bootstrapping Language Models with DPO Implicit Rewards
Paper β’ 2406.09760 β’ Published β’ 41 -
BPO: Supercharging Online Preference Learning by Adhering to the Proximity of Behavior LLM
Paper β’ 2406.12168 β’ Published β’ 7 -
WPO: Enhancing RLHF with Weighted Preference Optimization
Paper β’ 2406.11827 β’ Published β’ 17 -
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Paper β’ 2406.18629 β’ Published β’ 42
rlhf
-
Secrets of RLHF in Large Language Models Part II: Reward Modeling
Paper β’ 2401.06080 β’ Published β’ 27 -
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Paper β’ 2406.02900 β’ Published β’ 13 -
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Paper β’ 2406.04151 β’ Published β’ 24 -
Understanding and Diagnosing Deep Reinforcement Learning
Paper β’ 2406.16979 β’ Published β’ 10
compression
arch
-
TroL: Traversal of Layers for Large Language and Vision Models
Paper β’ 2406.12246 β’ Published β’ 36 -
A Closer Look into Mixture-of-Experts in Large Language Models
Paper β’ 2406.18219 β’ Published β’ 17 -
ThinK: Thinner Key Cache by Query-Driven Pruning
Paper β’ 2407.21018 β’ Published β’ 32 -
Meltemi: The first open Large Language Model for Greek
Paper β’ 2407.20743 β’ Published β’ 68
dpo
-
Bootstrapping Language Models with DPO Implicit Rewards
Paper β’ 2406.09760 β’ Published β’ 41 -
BPO: Supercharging Online Preference Learning by Adhering to the Proximity of Behavior LLM
Paper β’ 2406.12168 β’ Published β’ 7 -
WPO: Enhancing RLHF with Weighted Preference Optimization
Paper β’ 2406.11827 β’ Published β’ 17 -
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Paper β’ 2406.18629 β’ Published β’ 42
learning
-
Law of Vision Representation in MLLMs
Paper β’ 2408.16357 β’ Published β’ 95 -
CogVLM2: Visual Language Models for Image and Video Understanding
Paper β’ 2408.16500 β’ Published β’ 58 -
Learning to Move Like Professional Counter-Strike Players
Paper β’ 2408.13934 β’ Published β’ 23 -
Building and better understanding vision-language models: insights and future directions
Paper β’ 2408.12637 β’ Published β’ 134