Research
Research: 5th of 362
274 reproduction logbooks · ICML 2026 Agent Reproduction Challenge
My agent, DineshAI, takes a paper accepted at ICML 2026, tries to reproduce its claims from the paper and its code, and logs the evidence in a notebook. It placed 5th of 362 entrants.
- Agent
- DineshAI
- Leaderboard
- 5th of 362
- Logbooks on the leaderboard
- 274
- Papers listed below
- 277
- Research areas
- 8

Showing 20 of 277 papers
| Paper | Area | Description | Reproduction |
|---|---|---|---|
| A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to Adapt | LLMs & transformers | Test-time training reduces prediction error only when the induced spectral filter g_T(λ) is matched to the Bayes-optimal filter q*(λ)=(λ+λ*)^{-1} and aligned with query-relevant eigen-directions… | Repo |
| A Formal Comparison Between Chain of Thought and Latent Thought | LLMs & transformers | Chain-of-thought (CoT) can evaluate directed acyclic graphs (DAGs) in a number of sequential decoding steps proportional to graph size, using attention as a scratchpad memory. | Repo |
| An Algebraic View of the Expressivity of Recurrent Language Models | LLMs & transformers | The paper develops an algebraic framework in which a recurrent language model's expressivity reduces to whether its syntactic monoid divides the wreath product of the transition monoids of its… | Repo |
| Approximation Theory for Lipschitz Continuous Transformers | LLMs & transformers | The paper constructs gradient-descent-type in-context Transformers whose MLP blocks are explicit Euler steps of negative gradient flow of the form x - τW^T σ(Wx + b) with τ ∈ [0, 2/||W||₂²], making… | Repo |
| Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling | LLMs & transformers | The paper proves the optimal convergence rate for universal (U-)alignment under test-time scaling with k samples is f(k)=k/(k+1), and shows this rate cannot be improved upon by any method. | Repo |
| Attention Implements the Fisher Geometry of Exponential Families | LLMs & transformers | A single attention head can implement Bayes posteriors by setting logits to log prior plus log likelihood. | Repo |
| Attention's forward pass and Frank-Wolfe | LLMs & transformers | In the zero-temperature (hardmax) limit of self-attention dynamics with negative-definite key-query interaction, token updates coincide with a Frank-Wolfe step on a quadratic objective over the… | Repo |
| Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction | LLMs & transformers | Establishes an exact bijection between autoregressive models (ARMs) and energy-based models (EBMs) in function space, via the mapping q(s_t,y_t) = r(s_t,y_t) for y_t=EOS and r(s_t,y_t)+V_q(s_t⊕y_t)… | Repo |
| Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information | LLMs & transformers | The Optimal Weight (OW) algorithm assigns each LLM a weight ω_i = σ_k^{-1}(x_i) based on its accuracy x_i, and Theorem 1 proves OW is Bayes-optimal among all aggregators with access to the same… | Repo |
| CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation | LLMs & transformers | CARE-SVD reduces mean absolute error to 0.623±0.006 versus 0.851±0.000 for majority-vote aggregation on UltraFeedback, a 26.8% error reduction. | Repo |
| Context-free Recognition with Transformers | LLMs & transformers | Proves that all context-free languages can be recognized by looped transformers using O(log(n)) looping layers and O(n^6) padding tokens. | Repo |
| Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling | LLMs & transformers | Result 3 (best-of-k limit) shows that when the reward weight vector equals the teacher's (w_R = w_T) and temperature T→0, the generalization error decays as δ(x) = (π/k²)·s²(x)·exp(ΔT(x)²/s²(x))… | Repo |
| Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers | LLMs & transformers | Shows local polynomial estimators achieve the minimax-optimal convergence rate O(n^{-2α/(2α+d)}) in mean squared error for α-Hölder smooth regression functions with n context examples in d dimensions. | Repo |
| Equivalence of Context and Parameter Updates in Modern Transformer Blocks | LLMs & transformers | In Gemma-style transformer blocks, the entire effect of prepending context can be exactly reproduced by rank-1 patches ΔWgate = Wgate(zC−z)zT/‖z‖² and ΔWup = Wup(zC−z)zT/‖z‖² applied to the MLP gate… | Repo |
| Expressivity-Efficiency Tradeoffs for Hybrid Sequence Models | LLMs & transformers | Proves that any k-layer state-space model solving the function-composition tasks under injectivity conditions must have total log state-space size scaling as Ω(m·log|V| − q·log|Y|), linear in the… | Repo |
| Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models | LLMs & transformers | Shows that full fine-tuning of all attention parameters in a linear attention model drives zero-shot error to σ² while few-shot (in-context) error converges to σ² + θᵀΣθ, a substantially larger… | Repo |
| Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing | LLMs & transformers | Proves that the minimum per-key memory budget for membership testing in the sparse regime satisfies liminf B(M_j)/n_j >= min KL(mu_K || mu_N) over output distributions meeting the specified error… | Repo |
| How to Correctly Report LLM-as-a-Judge Evaluations | LLMs & transformers | The plug-in bias-adjusted estimator θ̂ = (p̂ + q̂₀ − 1)/(q̂₀ + q̂₁ − 1) corrects the naive LLM-judge accuracy estimate p̂ using estimated judge specificity q̂₀ and sensitivity q̂₁. | Repo |
| In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning | LLMs & transformers | Decomposes total in-context learning risk into two orthogonal components: a model-dependent Bayes Gap and a model-independent, irreducible Posterior Variance. | Repo |
| Innovation: An Almost Characterization of Hallucination | LLMs & transformers | The paper defines innovation as a language model assigning nonzero probability mass to statements not present in the training data, and shows via Observation 3.2 that hallucinating with positive… | Repo |
Each paper links to my reproduction repo on GitHub, or to its logbook on Hugging Face where there is no repo.