Why Muon Outperforms Adam: A Curvature Perspective (中文翻译) Permalink
Published:
A translation and walk-through of the curvature argument for why the Muon optimizer beats Adam — what the loss-landscape geometry says about each update rule.
No matches — try another term.
Published:
A translation and walk-through of the curvature argument for why the Muon optimizer beats Adam — what the loss-landscape geometry says about each update rule.
Published:
A Chinese translation of a paper that learns binary sampling patterns for single-pixel imaging via bilevel optimisation — jointly optimising the sensing pattern and the reconstruction network.
Published:
A Chinese translation of the PathCTM paper — a continuous-thought model that adapts its reasoning depth to accelerate analysis of gigapixel pathology images.
Published:
A full Chinese translation and close reading of the LoRA paper (Hu et al., 2021) — formulas rendered with MathJax, the original vector figures inlined, and tables rebuilt from the source.
Published:
A side-by-side walk-through of two vision-language architectures — how LLaVA bolts a vision encoder onto an LLM versus how Qwen2-VL fuses modalities natively — with diagrams of each fusion path.
Published:
A Chinese translation of “A guide to stochastic optimisation for large-scale inverse problems” — how SGD-family methods scale to imaging inverse problems, with the accompanying analysis and experiments.
Published:
A full Chinese translation and close reading of the DPO paper — how preference alignment can skip the reward model and RL loop, with the derivation and experiments reproduced from the original.
Published:
A read-through and Chinese translation of the diffusion language models survey — how denoising-diffusion ideas carry over from images to discrete text generation.
Published:
Notes and translation on Cola DLM — running diffusion language modeling in a continuous latent space instead of over discrete tokens.