<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Welcome to Jikai Jin's website</title><link>https://jkjin.com/</link><description>Recent content on Welcome to Jikai Jin's website</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 02 Aug 2026 11:21:23 -0400</lastBuildDate><atom:link href="https://jkjin.com/index.xml" rel="self" type="application/rss+xml"/><item><title>Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning</title><link>https://jkjin.com/publication/llm_causal_representation_learning/</link><pubDate>Wed, 08 Jul 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/llm_causal_representation_learning/</guid><description>We propose a causal representation learning framework to understand the hierarchical structure of language model capabilities, revealing causal relationships between general problem-solving, instruction-following, and mathematical reasoning abilities.</description></item><item><title>Salt Lake City</title><link>https://jkjin.com/gallery/salt_lake_city/</link><pubDate>Wed, 27 May 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/salt_lake_city/</guid><description/></item><item><title>A Long Night in the Farm</title><link>https://jkjin.com/gallery/farm_night/</link><pubDate>Fri, 24 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/farm_night/</guid><description/></item><item><title>A Visit to Rwanda</title><link>https://jkjin.com/gallery/rwanda/</link><pubDate>Fri, 24 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/rwanda/</guid><description/></item><item><title>Enchanted Forest of Light</title><link>https://jkjin.com/gallery/descanso_garden/</link><pubDate>Fri, 24 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/descanso_garden/</guid><description/></item><item><title>The Capital of New England</title><link>https://jkjin.com/gallery/boston/</link><pubDate>Fri, 24 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/boston/</guid><description/></item><item><title>The City of Datong</title><link>https://jkjin.com/gallery/datong/</link><pubDate>Fri, 24 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/datong/</guid><description/></item><item><title>The City That Never Sleeps</title><link>https://jkjin.com/gallery/new_york/</link><pubDate>Fri, 24 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/new_york/</guid><description/></item><item><title>Spring at Stanford</title><link>https://jkjin.com/gallery/stanford_spring/</link><pubDate>Thu, 23 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/stanford_spring/</guid><description/></item><item><title>Sunny Seattle</title><link>https://jkjin.com/gallery/seattle/</link><pubDate>Mon, 20 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/seattle/</guid><description/></item><item><title>Just Another Ocean-side Town</title><link>https://jkjin.com/gallery/oceanside/</link><pubDate>Wed, 15 Apr 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/gallery/oceanside/</guid><description/></item><item><title>Prescriptive Scaling Reveals the Evolution of Language Model Capabilities</title><link>https://jkjin.com/publication/prescriptive-scaling/</link><pubDate>Tue, 17 Feb 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/prescriptive-scaling/</guid><description>Prescriptive scaling laws: estimate stable capability boundaries vs. pretraining compute, track temporal shifts, and release the Proteus 2k evaluation dataset.</description></item><item><title>Prescriptive Scaling for Language Models</title><link>https://jkjin.com/post/prescriptive-scaling/</link><pubDate>Sat, 07 Feb 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/post/prescriptive-scaling/</guid><description>Given a pre-training compute budget, we can predict the best downstream performance strong post-training can reliably attain, and detect when that boundary shifts.</description></item><item><title>Adaptive Exploration for Latent-State Bandits</title><link>https://jkjin.com/publication/latent-bandit/</link><pubDate>Wed, 04 Feb 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/latent-bandit/</guid><description>We propose adaptive, state-model-free bandit algorithms for latent-state (confounded, non-stationary) environments.</description></item><item><title>Policy Learning with Abstention</title><link>https://jkjin.com/publication/abstention/</link><pubDate>Wed, 28 Jan 2026 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/abstention/</guid><description>We study policy learning with the ability to abstain, and demonstrate its usefulness in ensuring safety.</description></item><item><title>Sharp Structure-Agnostic Lower Bounds for General Linear Functional Estimation</title><link>https://jkjin.com/publication/general-functional/</link><pubDate>Fri, 05 Dec 2025 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/general-functional/</guid><description>Sharp minimax theory for structure-agnostic estimation of general linear functionals, including optimality of first-order debiasing and doubly robust estimators.</description></item><item><title>It's Hard to Be Normal: The Impact of Noise on Structure-agnostic Estimation</title><link>https://jkjin.com/publication/hard_to_be_normal/</link><pubDate>Thu, 03 Jul 2025 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/hard_to_be_normal/</guid><description>We consider structure-agnostic causal inference that estimates treatment effect estimation using black-box ML estimates of nuisance functions, and show that the celebrated DML is optimal when the treatment noise is Gaussian. When the noise is non-Gaussian, we propose ACE, a novel class of higher-order structure-agnostic estimators.</description></item><item><title>Solving Inequality Proofs with Large Language Models</title><link>https://jkjin.com/publication/llm_inequality/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/llm_inequality/</guid><description>We build a novel benchmark for evaluating how well can state-of-the-art LLMs solve inequality proving problems. Our evaluation studies provide insights on their advanced math reasoning capabilities and highlight their common reasoning flaws.</description></item><item><title>Structure-agnostic Optimality of Doubly Robust Learning for Treatment Effect Estimation</title><link>https://jkjin.com/publication/optimality-of-debiasing-for-ate-estimation/</link><pubDate>Fri, 02 May 2025 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/optimality-of-debiasing-for-ate-estimation/</guid><description>We show that first-order debiasing of black-box ML estimators is optimal for estimating average treatment effect.</description></item><item><title>Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking</title><link>https://jkjin.com/publication/grokking/</link><pubDate>Sat, 30 Nov 2024 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/grokking/</guid><description>We investigate the &amp;ldquo;grokking&amp;rdquo; phenomenon in deep learning on some simple setups, and show that it is caused by a dichotomy of the implicit biases between the early phase and late phase during training.</description></item><item><title>Learning Causal Representations from General Environments: Identifiability and Intrinsic Ambiguity</title><link>https://jkjin.com/publication/causal-representation-learning/</link><pubDate>Wed, 25 Sep 2024 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/causal-representation-learning/</guid><description>We study the best-achievable identification guarantees and provable identification algorithms for causal representation learning when hard interventions are not available.</description></item><item><title>Minimax Optimal Kernel Operator Learning via Multilevel Training</title><link>https://jkjin.com/publication/operlearning/</link><pubDate>Sat, 30 Sep 2023 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/operlearning/</guid><description>We consider the problem of learning a linear operator between Sobolev RKHSs from noisy data. Different from its finite-dimensional counterpart where regularized least squares is optimal, we prove that estimators with a certain multilevel structure is necessary (and sufficient) to achieve optimality.</description></item><item><title>Understanding Incremental Learning of Gradient Descent -- A Fine-grained Analysis of Matrix Sensing</title><link>https://jkjin.com/publication/matrixsensing/</link><pubDate>Fri, 27 Jan 2023 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/matrixsensing/</guid><description>We prove that GD applied to the matrix sensing problem has intriguing properties &amp;ndash; with small initialization and early stopping, it follows an incremental/greedy low-rank learning procedure. This form of simplicity bias allows GD to recover the ground-truth, despite over-parameterization and non-convexity.</description></item><item><title>Why Robust Generalization in Deep Learning is Difficult: Perspective of Expressive Power</title><link>https://jkjin.com/publication/robustness/</link><pubDate>Fri, 27 May 2022 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/robustness/</guid><description>We provide theoretical evidence that the hardness of robust generalization may stem from the expressive power of deep neural networks. Even when standard generalization is easy, robust generalization provably requires the size of DNNs to be exponentially large.</description></item><item><title>Understanding Riemannian Acceleration via a Proximal Extragradient Framework</title><link>https://jkjin.com/publication/riemann_ahpe/</link><pubDate>Thu, 10 Feb 2022 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/riemann_ahpe/</guid><description>We provide an improved analysis of the convergence rates of clipping algorithms, theoretically justifying their superior performance in deep learning.</description></item><item><title>Non-convex Distributionally Robust Optimization: Non-asymptotic Analysis</title><link>https://jkjin.com/publication/dro/</link><pubDate>Tue, 05 Oct 2021 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/dro/</guid><description>We proposed the first non-asymptotic analysis of algorithms for DRO with non-convex losses. Our algorithm incorporates momentum and adaptive step size, and has superior empirical performance.</description></item><item><title>Improved analysis of clipping algorithms for non-convex optimization</title><link>https://jkjin.com/publication/example/</link><pubDate>Mon, 05 Oct 2020 00:00:00 +0000</pubDate><guid>https://jkjin.com/publication/example/</guid><description>We provide an improved analysis of the convergence rates of clipping algorithms, theoretically justifying their superior performance in deep learning.</description></item><item><title>Awards and Honors</title><link>https://jkjin.com/awards/</link><pubDate>Fri, 01 Dec 2017 00:00:00 +0000</pubDate><guid>https://jkjin.com/awards/</guid><description>List of awards and honors.</description></item><item><title>News</title><link>https://jkjin.com/news/</link><pubDate>Fri, 01 Dec 2017 00:00:00 +0000</pubDate><guid>https://jkjin.com/news/</guid><description>List of news.</description></item><item><title/><link>https://jkjin.com/admin/config.yml</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jkjin.com/admin/config.yml</guid><description/></item></channel></rss>