<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>混沌有序</title><link>https://blog-seeker.pages.dev</link><description>于混沌淘金，向秩序前行。致模型，亦致求索者。</description><language>zh-CN</language><lastBuildDate>Wed, 13 May 2026 00:00:00 GMT</lastBuildDate><atom:link href="https://blog-seeker.pages.dev/feed.xml" rel="self" type="application/rss+xml" /><item><title>显存刺客 KV Cache：从算清账到撑起百万上下文</title><link>https://blog-seeker.pages.dev/posts/kv_cache</link><guid isPermaLink="true">https://blog-seeker.pages.dev/posts/kv_cache</guid><pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate><description>KV Cache 已经不只是一个推理优化细节，而是长上下文、多模态、Agent 服务的成本中枢。本文从显存公式讲起，沿系统管理、模型架构、动态压缩、量化与 Offloading 五条路线，重新整理 2026 年前后的 KV Cache 技术版图。</description><dc:creator>AI 推理优化研究</dc:creator><category>KV Cache</category><category>LLM</category><category>推理优化</category><category>长上下文</category><category>系统架构</category></item><item><title>万字长文：LLM 高阶解码策略与推理加速范式全景深度解析</title><link>https://blog-seeker.pages.dev/posts/advanced_decoding</link><guid isPermaLink="true">https://blog-seeker.pages.dev/posts/advanced_decoding</guid><pubDate>Thu, 05 Mar 2026 00:00:00 GMT</pubDate><description>从概率分布干预到系统级软硬件协同，深度拆解 DoLa、CFG、投机解码家族 (EAGLE/SSD) 以及多词预测 (Medusa/MTPC) 的数学原理与工程实现。</description><dc:creator>混沌有序</dc:creator><category>LLM</category><category>Inference</category><category>Speculative Decoding</category><category>MTP</category><category>System 2</category><category>Reinforcement Learning</category></item><item><title>后 GRPO 时代：长 CoT 的爆发与“对齐”新挑战</title><link>https://blog-seeker.pages.dev/posts/grpo_varieties</link><guid isPermaLink="true">https://blog-seeker.pages.dev/posts/grpo_varieties</guid><pubDate>Wed, 25 Feb 2026 00:00:00 GMT</pubDate><description>从 DAPO、GSPO、LUSPO 到 Dr.GRPO、GMPO、PMPO，一文拆解 GRPO 变体的核心动机、目标函数与工程改造路径。</description><dc:creator>混沌有序</dc:creator><category>RLHF</category><category>GRPO</category><category>Post-training</category></item><item><title>资源受限下的破局者：On-Policy Distillation</title><link>https://blog-seeker.pages.dev/posts/opd</link><guid isPermaLink="true">https://blog-seeker.pages.dev/posts/opd</guid><pubDate>Wed, 25 Feb 2026 00:00:00 GMT</pubDate><description>在算力与标注数据双重受限下，On-Policy Distillation 如何用 Reverse KL、γ=0 与在线探索实现高效后训练。</description><dc:creator>混沌有序</dc:creator><category>LLM</category><category>Distillation</category><category>Post-training</category></item><item><title>大模型字斟句酌的暗箱操作：Decoding 算法全景硬核拆解</title><link>https://blog-seeker.pages.dev/posts/decoding_algorithms</link><guid isPermaLink="true">https://blog-seeker.pages.dev/posts/decoding_algorithms</guid><pubDate>Mon, 23 Feb 2026 00:00:00 GMT</pubDate><description>Greedy、Beam、Top-K、Top-P、Min-P、Contrastive 与 Constrained Decoding 的底层机制与 PyTorch 实现全景拆解。</description><dc:creator>混沌有序</dc:creator><category>LLM</category><category>Decoding</category><category>Sampling</category></item><item><title>参数高效微调的终极炼丹术：从 LoRA 到它的“魔改”宇宙</title><link>https://blog-seeker.pages.dev/posts/lora</link><guid isPermaLink="true">https://blog-seeker.pages.dev/posts/lora</guid><pubDate>Sun, 22 Feb 2026 00:00:00 GMT</pubDate><description>从 LoRA、LoRA+、PiSSA、DoRA 到 TinyLoRA 与 LoRA-Mixer，一文看懂参数高效微调的核心机制与工程取舍。</description><dc:creator>混沌有序</dc:creator><category>LoRA</category><category>PEFT</category><category>Fine-tuning</category></item><item><title>后 DPO 时代的百家争鸣：如何优雅地给大模型“立规矩“</title><link>https://blog-seeker.pages.dev/posts/dpo_varieties</link><guid isPermaLink="true">https://blog-seeker.pages.dev/posts/dpo_varieties</guid><pubDate>Sat, 21 Feb 2026 00:00:00 GMT</pubDate><description>各大顶级实验室爆改 DPO出了一套极其华丽的招式表</description><dc:creator>混沌有序</dc:creator><category>Reinforcement Learning</category></item><item><title>人类偏好的刻度：PPO、DPO 与 GRPO 极简拆解</title><link>https://blog-seeker.pages.dev/posts/ppo_dpo_grpo</link><guid isPermaLink="true">https://blog-seeker.pages.dev/posts/ppo_dpo_grpo</guid><pubDate>Fri, 20 Feb 2026 00:00:00 GMT</pubDate><description>在探讨大语言模型（LLM）的对齐算法之前，我们需要先理清模型训练的宏观图景。大模型的训练通常分为三个阶段：预训练（Pre-training）、指令微调（SFT）和人类偏好对齐（RLHF/Alignment）。</description><dc:creator>混沌有序</dc:creator><category>Reinforcement Learning</category></item></channel></rss>