Model Mechanics
2026
10
- What Should World Models Be, and How Should We Use Them?
- JoyAI-VL-Interaction: From Chat Back to Continuous Interaction
- Bad Is Good: Why DeepSeek Did Not Use an n-Gram Structure
- Reading Model States Through Newline Tokens: What Word Salad Chopper Reveals
- From LSH to K-Center Greedy: Semantic Embeddings for Deduplication, Cleaning, and Sample Selection
- Why Language Models Hallucinate
- What Does the Loss Landscape of LLMs Look Like?
- Compression for AGI: Compression as Intelligence
- Neural Scaling Laws: From Kaplan to Chinchilla
- From LLM to VLM: How Language Models Learn Visual Understanding