-
LoRA: Low-Rank Adaptation of Large Language Models
Paper • 2106.09685 • Published • 64 -
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Paper • 2205.14135 • Published • 15 -
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Paper • 2201.11903 • Published • 15 -
Training language models to follow instructions with human feedback
Paper • 2203.02155 • Published • 25
Vismay Patel
vismayapatel
AI & ML interests
None yet
Organizations
None yet
NLP basics
-
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Paper • 2006.03654 • Published • 3 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 32 -
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Paper • 1907.11692 • Published • 10 -
Attention Is All You Need
Paper • 1706.03762 • Published • 133
VLM Foundational
-
LoRA: Low-Rank Adaptation of Large Language Models
Paper • 2106.09685 • Published • 64 -
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Paper • 2205.14135 • Published • 15 -
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Paper • 2201.11903 • Published • 15 -
Training language models to follow instructions with human feedback
Paper • 2203.02155 • Published • 25
NLP basics
-
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Paper • 2006.03654 • Published • 3 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 32 -
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Paper • 1907.11692 • Published • 10 -
Attention Is All You Need
Paper • 1706.03762 • Published • 133
models 0
None public yet
datasets 0
None public yet