Foundation Models
2026
10
- Can Closed-Source Models Be Distilled? Knowledge Distillation for Generative Language Models
- What Should World Models Be, and How Should We Use Them?
- JoyAI-VL-Interaction: From Chat Back to Continuous Interaction
- Bad Is Good: Why DeepSeek Did Not Use an n-Gram Structure
- Reading Model States Through Newline Tokens: What Word Salad Chopper Reveals
- Reward Hacking: When Optimizers Reverse-Search the Reward Signal
- From LSH to K-Center Greedy: Semantic Embeddings for Deduplication, Cleaning, and Sample Selection The Evolution of Reward Design: From RLHF to RLVR
- Reinforcement Learning in LLM Alignment: From Reward Signals to Advantage Estimation
- Parameter-Efficient Fine-Tuning (PEFT): From Adapter to LoRA