Integrating Inductive Biases in Transformers via Distillation for Financial Time Series Forecasting Paper • 2603.16985 • Published Mar 17 • 2
view article Article How to Use Multiple GPUs in Hugging Face Transformers: Device Map vs Tensor Parallelism ariG23498 • Feb 12 • 20
Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning Paper • 2602.04998 • Published Feb 4 • 7
Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning Paper • 2602.04998 • Published Feb 4 • 7
Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning Paper • 2602.04998 • Published Feb 4 • 7
sentence-transformers/all-mpnet-base-v2 Sentence Similarity • 0.1B • Updated Aug 19, 2025 • 27.5M • • 1.34k
Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions Paper • 2506.08234 • Published Jun 9, 2025 • 9
Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions Paper • 2506.08234 • Published Jun 9, 2025 • 9
Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions Paper • 2506.08234 • Published Jun 9, 2025 • 9 • 3