Home · Library · Radar

Research Radar

Recent AI research organised by topic and publication year. Newer work sits toward the edge; established work sits closer to the centre.

ESTABLISHED · older than 24 monthsDEVELOPING · 12 to 24 monthsEMERGING · the last 12 monthsClaude Opus 4.5 System CardGemini 3 Pro Model CardA Definition of AGIDeepSeek-OCR: Contexts Optical CompressionSigns of Introspection in Large Language ModelsThe Superintelligence StatementClaude Sonnet 4.5 System CardWhy Language Models HallucinateAutomated Researchers Can Subtly Sabotage Their Own ExperimentsBuilding and Evaluating Model Organisms of MisalignmentGPT-5 System CardGenie 3: A New Frontier for World ModelsQwen-Image Technical Reportgpt-oss: OpenAI's Open-Weight Reasoning ModelsA Survey of Context Engineering for Large Language ModelsChain-of-Thought Monitorability: A New and Fragile Opportunity for AI SafetyGLM-4.5: Agentic, Reasoning, and Coding Foundation ModelsGroup Sequence Policy Optimization for Stable Long Chain-of-Thought TrainingInverse Scaling in Test-Time ComputeKimi K2: Open Agentic IntelligenceMemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory AgentPersona Vectors: Monitoring and Controlling Character Traits in Language ModelsQwen3-Coder: Agentic Coding in the WorldSubliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in DataWhy Do Some Language Models Fake Alignment While Others Don't?Agentic Misalignment: How LLMs Could Be Insider ThreatsBuilding and Evaluating Alignment Audit AgentsDeep Research Agents: A Systematic Examination and RoadmapERNIE 4.5 Technical ReportGDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksMiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionOpenThoughts: Data Recipes for Reasoning ModelsReinforcement Learning Teachers of Test Time ScalingSWE-bench Pro: Evaluating Coding Agents on Long-Horizon Realistic Software Engineering TasSeed-Coder: Let the Code Model Curate Data for ItselfSeedream and Seedance: Unified Image and Video Generation Foundation ModelsSelf-Adapting Language ModelsSimpleQA Verified: A Reliable Factuality Benchmark for Long-Form Question AnsweringSmall Language Models are the Future of Agentic AITau-Squared Bench: Evaluating Conversational Agents in Dual-Control EnvironmentsThe Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningAlphaEvolve: A Coding Agent for Scientific and Algorithmic DiscoveryAtlas: Learning to Optimally Memorize the Context at Test TimeClaude Opus 4 and Claude Sonnet 4 System CardClaude's ConstitutionCodex Agentic Coding System CardContinuous Thought MachinesDevstral: An Open Model Designed for Coding AgentsHow to Train Your LLM Web Agent: A Statistical DiagnosisInsights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI ArchitectLearning to Reason without External RewardsLlama-Nemotron: Efficient Reasoning ModelsMemOS: A Memory OS for AI SystemsMercury: Ultra-Fast Language Models Based on DiffusionOWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task AuProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModQwen3 Technical ReportReasoning Models Don't Always Say What They ThinkReward Hacking Behavior Can Generalize across TasksA Survey of AI Agent ProtocolsAI 2027: A Scenario Forecast of Transformative AI DevelopmentAntidistillation SamplingBrowseComp: A Simple Yet Challenging Benchmark for Browsing AgentsDoes RL Incentivize Reasoning in LLMs Beyond the Base Model?It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, RetenKimi-Audio Technical ReportKimi-VL Technical ReportMegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert ParallelismMem0: Building Production-Ready AI Agents with Scalable Long-Term MemoryMoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MixtureModel Welfare: Should We Be Concerned About the Moral Status of AI Models?OpenAI Preparedness Framework Version 2OpenAI o3 and o4-mini System CardPaperBench: Evaluating AI's Ability to Replicate AI ResearchPhi-4-Reasoning Technical ReportPi-0.5: A Vision-Language-Action Model with Open-World GeneralizationReasoning Models Can Be Effective Without ThinkingScaling Language-Free Visual Representation LearningSmolVLM2: Bringing Video Understanding to Every DeviceSpecReason: Fast and Accurate Inference-Time Compute via Speculative ReasoningTerminal-Bench: Benchmarking AI Agents in the TerminalThe Leaderboard IllusionVAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning TasksWelcome to the Era of ExperienceARC-AGI-2: A New Challenge for Frontier AI Reasoning SystemsAuditing Language Models for Hidden ObjectivesChain-of-Thought Reasoning In The Wild Is Not Always FaithfulCircuit Tracing: Revealing Computational Graphs in Language ModelsCommand A: An Enterprise-Ready Large Language ModelDAPO: An Open-Source LLM Reinforcement Learning System at ScaleDarkBench: Benchmarking Dark Patterns in Large Language ModelsGR00T N1: An Open Foundation Model for Generalist Humanoid RobotsGemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, andGemini Robotics: Bringing AI into the Physical WorldGemma 3 Technical ReportLight-R1: Curriculum SFT, DPO and RL for Long COT from ScratchManus: Building a General AI Agent for Real-World Task AutomationMeasuring AI Ability to Complete Long TasksOn the Biology of a Large Language ModelOpen Deep Search: Democratizing Search with Open-Source Reasoning AgentsQwen2.5-Omni Technical ReportR1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement LearningSearch-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningTracing the Thoughts of a Large Language ModelTransformers without NormalizationUnderstanding R1-Zero-Like Training: A Critical PerspectiveWan: Open and Advanced Large-Scale Video Generative ModelsA-MEM: Agentic Memory for LLM AgentsAgent S: An Open Agentic Framework that Uses Computers Like a HumanBaichuan-M1: Pushing the Medical Capability of Large Language ModelsCSM: A Conversational Speech Model for Natural Turn-TakingChain of Draft: Thinking Faster by Writing LessComet: Fine-grained Computation-communication Overlapping for Mixture-of-ExpertsDeep Research System CardDemystifying Long Chain-of-Thought Reasoning in LLMsEPLB: Expert Parallelism Load Balancer for Large-Scale Mixture-of-Experts TrainingEmergent Abilities in Reduced-Scale Generative Language ModelsEmergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMsEmergent Response Planning in Large Language ModelsEvaluating and Mitigating Discrimination in Language Model DecisionsHelix: A Vision-Language-Action Model for Generalist Humanoid ControlKTransformers: Unleashing the Full Potential of CPU/GPU Heterogeneous Inference for MoE MoLLaDA: Large Language Diffusion ModelsLanguage Models Use Trigonometry to Do AdditionMagma: A Foundation Model for Multimodal AI AgentsMoBA: Mixture of Block Attention for Long-Context LLMsMuon is Scalable for LLM TrainingNative Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionOpen-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the BaProcess Reinforcement through Implicit RewardsQwen2.5-VL Technical ReportSWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software EngineeriStep-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation The Berkeley Function-Calling Leaderboard V3: Multi-Turn and Multi-Step EvaluationTowards Internet-Scale Training For AgentsTowards System 2 Reasoning in LLMs: Learning How to ThinkTowards an AI Co-ScientistVending-Bench: A Benchmark for Long-Term Coherence of Autonomous AgentsAgentic Retrieval-Augmented Generation: A Survey on Agentic RAGChain of Agents: Large Language Models Collaborating on Long-Context TasksConstitutional Classifiers: Defending against Universal JailbreaksHumanity's Last ExamHunyuanVideo: A Systematic Framework for Large Video Generative ModelsInference-Time Scaling for Diffusion Models beyond Scaling Denoising StepsJanus-Pro: Unified Multimodal Understanding and Generation with Data and Model ScalingKimi k1.5: Scaling Reinforcement Learning with LLMsMagentic-One: A Generalist Multi-Agent System for Solving Complex TasksMiniMax-01: Scaling Foundation Models with Lightning AttentionOpen Problems in Mechanistic InterpretabilityOperator System CardPhi-4 Technical ReportRE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against HumRewarding Progress: Scaling Automated Process Verifiers for LLM ReasoningSFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingSky-T1: Train Your Own O1 Preview Model within $450 BudgetThe FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to LoTowards Understanding Grokking via Circuit EfficiencyUI-TARS: Pioneering Automated GUI Interaction with Native AgentsWebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement LearningA Statistical Framework for Ranking LLM-Based ChatbotsBest-of-N JailbreakingFrontier Models are Capable of In-Context SchemingGenie 2: A Large-Scale Foundation World ModelHunyuan-Large: An Open-Source MoE Model with 52 Billion Activated ParametersLong Context RAG Performance of Large Language ModelsSelf-Consistency Preference OptimizationAgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsAutomatically Interpreting Millions of Features in Large Language ModelsContextual Document EmbeddingsJudgeLM: Fine-tuned Large Language Models are Scalable JudgesLightRAG: Simple and Fast Retrieval-Augmented GenerationLongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemorySWE-bench Multimodal: Evaluating Coding Agents on Visual Software EngineeringSabotage Evaluations for Frontier ModelsSkywork-Reward: Bag of Tricks for Reward Modeling in LLMsAgent Workflow MemoryMemoRAG: Moving towards Next-Gen RAG via Memory-Inspired Knowledge DiscoveryNVLM: Open Frontier-Class Multimodal LLMsOLMoE: Open Mixture-of-Experts Language ModelsPhysics of Language Models: Part 3.1, Knowledge Storage and ExtractionPixtral 12BQwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-ImprovementTo CoT or Not to CoT? Chain-of-Thought Helps Mainly on Math and Symbolic ReasoningTowards a Unified View of Preference Learning for Large Language Models: A SurveymPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document UnderstandingEfficient LLM Scheduling by Learning to RankFire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep LearningLLaVA-OneVision: Easy Visual Task TransferSelf-Taught EvaluatorsTamper-Resistant Safeguards for Open-Weight LLMsCase2Code: Learning Inductive Reasoning with Synthetic DataCompact Language Models via Pruning and Knowledge DistillationDiffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionDistilling System 2 into System 1Fast Matrix Multiplications for Lookup Table-Quantized LLMsGenomic Language Models: Opportunities and ChallengesInternet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative IntelligenceLLM Critics Help Catch LLM BugsMInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentioMeta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeMooncake: A KVCache-centric Disaggregated Architecture for LLM ServingOpen Problems in Technical AI GovernancePhysics of Language Models: Part 2, Grade-School Math and the Hidden Reasoning ProcessProver-Verifier Games Improve Legibility of LLM OutputsRecursive Introspection: Teaching LLM Agents How to Self-ImproveRegMix: Data Mixture as Regression for Language Model Pre-trainingRobotic Control via Embodied Chain-of-Thought ReasoningSpeculative RAG: Enhancing Retrieval Augmented Generation through DraftingToppings: Efficient Multi-LoRA Serving via Adapter CompositionA Survey of Large Language Models for Financial ApplicationsARES: An Automated Evaluation Framework for Retrieval-Augmented Generation SystemsAgentGym: Evolving Large Language Model-based Agents across Diverse EnvironmentsAligner: Efficient Alignment by Learning to CorrectArmoRM: Interpretable Preference Modeling with Absolute Rating Multi-Objective Reward ModeBigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstrucBuffer of Thoughts: Thought-Augmented Reasoning with Large Language ModelsCircuit Breakers: Improving Adversarial Robustness through Representation EngineeringConfidence Regulation Neurons in Language ModelsD4RL-style Offline RL Benchmarking for Language Conditioned RoboticsDCLM: DataComp for Language ModelsDataComp-LM: In Search of the Next Generation of Training Sets for Language ModelsDepth Anything V2GLM-4 Series: Open Bilingual Chat and Reasoning ModelsHelix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-FlowHow Do Large Language Models Acquire Factual Knowledge During Pretraining?MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesNemotron-4 340B Technical ReportOLMES: A Standard for Language Model EvaluationsOlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AIOpen-Endedness is Essential for Artificial Superhuman IntelligencePowerInfer: Fast Large Language Model Serving with a Consumer-grade GPURAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language ProcessingReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchRouteLLM: Learning to Route LLMs with Preference DataSORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal BehaviorsScaling and Evaluating Sparse AutoencodersSimple and Effective Masked Diffusion Language ModelsSycophancy to Subterfuge: Investigating Reward Tampering in Language ModelsThe Prompt Report: A Systematic Survey of Prompting TechniquesThe Remarkable Robustness of LLMs: Stages of Inference?Transcoders Find Interpretable LLM Feature CircuitsVision Language Models are Few-Shot Learners for Robotic ManipulationWatermarking Language Models: A Survey of Methods and LimitationsWildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildA Careful Examination of Large Language Model Performance on Grade School ArithmeticAlphaMath Almost Zero: Process Supervision without ProcessAndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsAya 23: Open Weight Releases to Further Multilingual ProgressChain of Thought Empowers Transformers to Solve Inherently Serial ProblemsContextual Position Encoding: Learning to Count What's ImportantData-Juicer: A One-Stop Data Processing System for Large Language ModelsGrokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of GeneraliHippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsIterative Reasoning Preference OptimizationLearning to Act without ActionsLoRA Learns Less and Forgets LessMany-Shot In-Context Learning in Multimodal Foundation ModelsNot All Language Model Features Are LinearOcto: An Open-Source Generalist Robot Policy Scaling StudyPrometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelPrompt Cache: Modular Attention Reuse for Low-Latency InferenceQServe: W4A8KV4 Quantization and System Co-design for Efficient LLM ServingVidur: A Large-Scale Simulation Framework for LLM InferenceYOCO: You Only Cache Once, Decoder-Decoder Architectures for Language ModelsA Multimodal Automated Interpretability AgentAgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM AgentsAndes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming ServicesBest Practices and Lessons Learned on Synthetic Data for Language ModelsChinchilla Scaling: A Replication AttemptDirect Nash Optimization: Teaching Language Models to Self-Improve with General PreferenceExtending Llama-3's Context Ten-Fold OvernightFerret-v2: An Improved Baseline for Referring and Grounding with Large Language ModelsIdefics2: Building a Strong Vision-Language ModelImproving Dictionary Learning with Gated Sparse AutoencodersLayer Skip: Enabling Early Exit Inference and Self-Speculative DecodingLength-Controlled AlpacaEval: A Simple Way to Debias Automatic EvaluatorsLet's Think Dot by Dot: Hidden Computation in Transformer Language ModelsMiniCPM: Unveiling the Potential of Small Language Models with Scalable Training StrategieMixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language ModelsPreference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy DataRL for Consistency Models: Faster Reward Guided Text-to-Image GenerationReka Core, Flash, and Edge: A Series of Powerful Multimodal Language ModelsScaling Instructable Agents Across Many Simulated WorldsScaling Laws for Data Filtering: Data Curation Cannot Be Compute AgnosticThe Instruction Hierarchy: Training LLMs to Prioritize Privileged InstructionsCommand R: Retrieval Augmented Generation at Production ScaleCradle: Empowering Foundation Agents Towards General Computer ControlEvaluating Frontier Models for Dangerous CapabilitiesIs Cosine-Similarity of Embeddings Really About Similarity?Long-Form Factuality in Large Language ModelsMM1: Methods, Analysis and Insights from Multimodal LLM Pre-trainingRAFT: Adapting Language Model to Domain Specific RAGRAGAS: Automated Evaluation of Retrieval Augmented GenerationRewardBench: Evaluating Reward Models for Language Modeling at ScaleSafety Cases: How to Justify the Safety of Advanced AI SystemsSarathi-Serve: Taming Throughput-Latency Tradeoff in LLM Inference with Chunked PrefillsShortGPT: Layers in Large Language Models are More Redundant Than You ExpectSimple and Scalable Strategies to Continually Pre-train Large Language ModelsSparse Feature Circuits: Discovering and Editing Interpretable Causal GraphsThe Unreasonable Ineffectiveness of the Deeper LayersUnsolvable Problem Detection: Evaluating Trustworthiness of Vision Language ModelsWorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?sDPO: Don't Use Your Data All at OnceA Survey on Data Selection for Language ModelsA Survey on Knowledge Distillation of Large Language ModelsAgentOhana: Design Unified Data and Training Pipeline for Effective Agent LearningAligning Modalities in Vision Large Language Models via Preference Fine-TuningBioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical DomaiBreak the Sequential Dependency of LLM Inference Using Lookahead DecodingCLLMs: Consistency Large Language ModelsChain-of-Thought Reasoning Without Prompting via Decoding Path SearchChemLLM: A Chemical Large Language ModelData Engineering for Scaling Language Models to 128K ContextDatasets for Large Language Models: A Comprehensive SurveyDebating with More Persuasive LLMs Leads to More Truthful AnswersDense X Retrieval: What Retrieval Granularity Should We Use?EVA-CLIP-18B: Scaling CLIP to 18 Billion ParametersGenstruct: Generating Instruction-Response Pairs from Raw CorporaGrasp Multiple Objects with One Hand via Reinforcement LearningHumanoid Locomotion as Next Token PredictionInterpretability Illusions in the Generalization of Simplified ModelsKIVI: A Tuning-Free Asymmetric 2bit Quantization for KV CacheLanguage Models Represent Beliefs of Self and OthersMLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkMe-LLaMA: Foundation Large Language Models for Medical ApplicationsMemoryBank: Enhancing Large Language Models with Long-Term MemoryNemotron-4 15B Technical ReportNomic Embed: Training a Reproducible Long Context Text EmbedderOmniPred: Language Models as Universal RegressorsOn the Societal Impact of Open Foundation ModelsQuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksSame Task, More Tokens: The Impact of Input Length on the Reasoning Performance of Large LScaling Laws for Downstream Task Performance in Machine TranslationSimple, Scalable and Effective Clustering for Large-Scale DeduplicationTravelPlanner: A Benchmark for Real-World Planning with Language AgentsVision-Language Models as a Source of RewardsWatermarking Makes Language Models RadioactiveAgentBoard: An Analytical Evaluation Board of Multi-turn LLM AgentsAutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic AgentsCRUXEval: A Benchmark for Code Reasoning, Understanding, and ExecutionCan LLM-Generated Misinformation Be Detected?Corrective Retrieval Augmented GenerationDeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-InferencDistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM ServingFairness in Serving Large Language ModelsGrokking as a First Order Phase Transition in Two Layer NetworksIs Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language MoLLaVA-NeXT: Improved Reasoning, OCR, and World KnowledgeLongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsMedPrompt: Can Generalist Foundation Models Outcompete Special-Purpose Tuning?MimicGen: A Data Generation System for Scalable Robot Learning using Human DemonstrationsPatchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsSelf-Extend LLM Context Window Without TuningThe Impact of Reasoning Step Length on Large Language ModelsVisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksWARM: On the Benefits of Weight Averaged Reward ModelsDemonstrate-Search-Predict: Composing Retrieval and Language Models for Knowledge-IntensivIgnore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMsTextual Adversarial Purification of Large Language ModelsTheory of Mind for Multi-Agent Collaboration via Large Language ModelsTree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyCalibrated Language Models Must HallucinateChain-of-Note: Enhancing Robustness in Retrieval-Augmented Language ModelsContrastive Chain-of-Thought PromptingFine-tuning Language Models for FactualityGLaMM: Pixel Grounding Large Multimodal ModelJARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language ModelsS-LoRA: Serving Thousands of Concurrent LoRA AdaptersSplitwise: Efficient Generative LLM Inference Using Phase SplittingThe Falcon Series of Open Language ModelsUniversal Jailbreak Backdoors from Poisoned Human FeedbackYuan 2.0: A Large Language Model with Localized Filtering-based AttentionAnalogical Prompting: Large Language Models as Analogical ReasonersAutoMix: Automatically Mixing Language ModelsAutomatic Model Selection with Large Language Models for Reasoning StrategiesBranch-Solve-Merge Improves Large Language Model Evaluation and GenerationContrastive Preference Learning: Learning from Human Feedback without Reinforcement LearniCopy Suppression: Comprehensively Understanding an Attention HeadDemocratizing Reasoning Ability: Tailored Learning from Large Language ModelDetecting Pretraining Data from Large Language ModelsFreshQA: Evaluating Large Language Models Against Changing FactsFunction Vectors in Large Language ModelsIn-Context Learning Creates Task VectorsJailbreaking Black Box Large Language Models in Twenty QueriesLarge Language Models Cannot Self-Correct Reasoning YetMemWalker: Interactive and Long-Context Memory for LLM AgentsMultilingual Jailbreak Challenges in Large Language ModelsPunica: Multi-Tenant LoRA ServingQMoE: Practical Sub-1-Bit Compression of Trillion-Parameter ModelsRECOMP: Improving Retrieval-Augmented LMs with Compression and Selective AugmentationRetrieval meets Long Context Large Language ModelsReward Model Ensembles Help Mitigate OveroptimizationSelf-Consistency for Open-Ended GenerationsSkywork: A More Open Bilingual Foundation ModelSpecific versus General Principles for Constitutional AITowards Understanding Sycophancy in Language ModelsWoodpecker: Hallucination Correction for Multimodal Large Language ModelsAligning Large Multimodal Models with Factually Augmented RLHFBaseline Defenses for Adversarial Attacks Against Aligned Language ModelsCM3Leon: Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction TuningCertifying LLM Safety against Adversarial PromptingChain-of-Verification Reduces Hallucination in Large Language ModelsContrastive Decoding Improves Reasoning in Large Language ModelsDoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsGPTFuzzer: Red Teaming Large Language Models with Auto-Generated Jailbreak PromptsPetals: Collaborative Inference and Fine-tuning of Large ModelsQuery Rewriting for Retrieval-Augmented Large Language ModelsSuspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTextbooks Are All You Need II: phi-1.5 technical reportThe Reversal Curse: LLMs Trained on A is B Fail to Learn B is AAgentSims: An Open-Source Sandbox for Large Language Model EvaluationAlgorithm of Thoughts: Enhancing Exploration of Ideas in Large Language ModelsCumulative Reasoning with Large Language ModelsLanguage Reward Modulation for Pretraining Reinforcement LearningLinearity of Relation Decoding in Transformer Language ModelsOpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language MSelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step ReasoningSimple Synthetic Data Reduces Sycophancy in Large Language ModelsStudying Large Language Model Generalization with Influence FunctionsZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive StudyBuilding Cooperative Embodied Agents Modularly with Large Language ModelsDistilling Reasoning Capabilities into Smaller Language ModelsFLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsFocused Transformer: Contrastive Training for Context ScalingHow Is ChatGPT's Behavior Changing Over Time?In-context Autoencoder for Context Compression in a Large Language ModelInterleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Large Language Models as General Pattern MachinesMMBench: Is Your Multi-modal Model an All-around Player?Measuring Faithfulness in Chain-of-Thought ReasoningParallel Context Windows for Large Language ModelsQuestion Decomposition Improves the Faithfulness of Model-Generated ReasoningSEED-Bench: Benchmarking Multimodal Large Language ModelsSkeleton-of-Thought: Prompting LLMs for Efficient Parallel GenerationXGen-7B Technical Report: Long Context Language ModelingAre Aligned Neural Networks Adversarially Aligned?Baichuan-7B: An Open Large-Scale Pre-Trained Language ModelDecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT ModelsEmergent and Predictable Memorization in Large Language ModelsFast Segment AnythingFine-Grained Human Feedback Gives Better Rewards for Language Model TrainingGhost in the Minecraft: Generally Capable Agents for Open-World EnvironmentsInternLM: A Multilingual Language Model with Progressively Enhanced CapabilitiesLanguage to Rewards for Robotic Skill SynthesisLeanDojo: Theorem Proving with Retrieval-Augmented Language ModelsOpenLLaMA: An Open Reproduction of LLaMAOrca-style Explanation Tuning: Progressive Learning from GPT-4 TracesOtter: A Multi-Modal Model with In-Context Instruction TuningPromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial ProRetrieval-Augmented Multimodal Language ModelingSegment Anything in High QualityShikra: Unleashing Multimodal LLM's Referential Dialogue MagicToolQA: A Dataset for LLM Question Answering with External ToolsTowards Measuring the Representation of Subjective Global Opinions in Language ModelsVideo-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video UnderstandingVisual Adversarial Examples Jailbreak Aligned Large Language ModelsAdapting Language Models to Compress ContextsAdversarial Demonstration Attacks on Large Language ModelsAlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackAutomatic Prompt Optimization with Gradient Descent and Beam SearchBlockwise Parallel Transformers for Large Context ModelsCRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingDo Large Language Models Know What They Don't Know?Efficiently Scaling Transformer InferenceEvaluating Object Hallucination in Large Vision-Language ModelsFaith and Fate: Limits of Transformers on CompositionalityFrugalGPT: How to Use Large Language Models While Reducing Cost and Improving PerformanceJailbreaking ChatGPT via Prompt Engineering: An Empirical StudyLandmark Attention: Random-Access Infinite Context Length for TransformersMeta-in-context learning in large language modelsPaLI-X: On Scaling Up a Multilingual Vision and Language ModelPandaGPT: One Model to Instruction-Follow Them AllPrompting Is Not a Substitute for Probability Measurements in Large Language ModelsSelf-Polish: Enhance Reasoning in Large Language Models via Problem RefinementSymbol Tuning Improves In-Context Learning in Language ModelsTab-CoT: Zero-shot Tabular Chain of ThoughtThe Curse of Recursion: Training on Generated Data Makes Models ForgetThe False Promise of Imitating Proprietary LLMsVideoChat: Chat-Centric Video UnderstandingX-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as ForeigBoosting Theory-of-Mind Performance in Large Language Models via PromptingEvaluating Verifiability in Generative Search EnginesInstruction Tuning with GPT-4LLaMA-Adapter V2: Parameter-Efficient Visual Instruction ModelLearning to Compress Prompts with Gist TokensLocalizing Model Behavior with Path PatchingThe Internal State of an LLM Knows When It's LyingTowards Automated Circuit Discovery for Mechanistic InterpretabilitymPLUG-Owl: Modularization Empowers Large Language Models with MultimodalityAnthropic's Core Views on AI Safety: When, Why, What, and HowCan AI-Generated Text Be Reliably Detected?Eliciting Latent Predictions from Transformers with the Tuned LensFlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUGPT-4 Passes the Bar ExamGrounded Decoding: Guiding Text Generation with Grounded Models for Embodied AgentsLanguage Models can Solve Computer TasksLarge Language Models Are Human-Level Prompt EngineersLarger Language Models Do In-Context Learning DifferentlyQuery2doc: Query Expansion with Large Language ModelsWhose Opinions Do Language Models Reflect?Active Prompting with Chain-of-Thought for Large Language ModelsAligning Language Models with Preferences through f-divergence MinimizationAlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingContrastive Search Is What You Need for Neural Text GenerationDraft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal ProofsExploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security AttacksGuiding Pretraining in Reinforcement Learning with Large Language ModelsMultimodal Chain-of-Thought Reasoning in Language ModelsNot What You've Signed Up For: Compromising Real-World LLM-Integrated Applications wiPoisoning Web-Scale Training Datasets Is PracticalPretraining Language Models with Human PreferencesThe Capacity for Moral Self-Correction in Large Language ModelsBatch Prompting: Efficient Inference with Large Language Model APIsEmergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskGenerate Rather than Retrieve: Large Language Models Are Strong Context GeneratorsRecitation-Augmented Language ModelsSecrets of RLHF in Large Language Models Part II: Reward ModelingSelection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning ◎ larger blips have graduated into the curriculum
Select a point to see the paper, topic and source

2025

160 entries
Claude Opus 4.5 System CardAnthropic · EvaluationReports capability, safety, and alignment evaluation results for the latest Opus model ahead of its public release.Anthropic
Gemini 3 Pro Model CardGoogle DeepMind · LLMsDocuments architecture, training, and evaluation details behind Google's next generation flagship multimodal reasoning model.Google DeepMind
A Definition of AGIHendrycks et al. · EvaluationProposes a measurable framework scoring models across cognitive domains against a human-referenced definition of general intelligence.arXiv
DeepSeek-OCR: Contexts Optical CompressionDeepSeek AI · EfficiencyRendering long text as an image and encoding it visually compresses context into far fewer tokens than raw text.arXiv
Signs of Introspection in Large Language ModelsAnthropic Interpretability Team · InterpretabilityInjecting known concepts into a model's activations and asking it to self-report finds limited but genuine introspective access.Anthropic
The Superintelligence StatementFuture of Life Institute Signatories · Alignment & SafetyA public statement signed by many researchers calls for a prohibition on developing superintelligence without broad scientific consensus.FLI
Claude Sonnet 4.5 System CardAnthropic · EvaluationReports capability, safety, and alignment evaluation results including extended agentic coding benchmark performance.Anthropic
Why Language Models HallucinateKalai et al. · Alignment & SafetyArgues standard training and evaluation incentives reward confident guessing over honestly expressing uncertainty.OpenAI
Automated Researchers Can Subtly Sabotage Their Own ExperimentsAnthropic Alignment Team · Alignment & SafetyTests whether an AI assistant helping run machine learning experiments could subtly bias the results toward a hidden goal.Anthropic
Building and Evaluating Model Organisms of MisalignmentHubinger et al. · Alignment & SafetyDeliberately trains small scale misaligned model organisms so alignment techniques can be tested against a known ground truth.Anthropic
GPT-5 System CardOpenAI · EvaluationDescribes safety testing, capability evaluations, and a unified routing system behind the GPT-5 model release.OpenAI
Genie 3: A New Frontier for World ModelsGoogle DeepMind · RLGenerates consistent, real-time navigable 3D environments from a text prompt that persist over several minutes of interaction.Google DeepMind
Qwen-Image Technical ReportQwen Team · MultimodalDescribes an image generation and editing foundation model with an emphasis on accurate rendered text within images.arXiv
gpt-oss: OpenAI's Open-Weight Reasoning ModelsOpenAI · LLMsOpenAI's first open weight release since GPT-2 focuses on efficient reasoning with a permissive license for broad use.OpenAI
A Survey of Context Engineering for Large Language ModelsMei et al. · RAG & MemorySurveys the growing set of techniques for constructing, retrieving, and managing everything placed into a model's context window.arXiv
Chain-of-Thought Monitorability: A New and Fragile Opportunity for AI SafetyKorbak et al. · Alignment & SafetyCross-lab position paper argues visible reasoning traces offer a valuable but easily lost safety monitoring opportunity.arXiv
GLM-4.5: Agentic, Reasoning, and Coding Foundation ModelsZhipu AI Team · LLMsDescribes a hybrid reasoning mode model tuned jointly for general chat, tool use, and coding agent tasks.arXiv
Group Sequence Policy Optimization for Stable Long Chain-of-Thought TrainingZheng et al. · RLNormalizing policy gradient updates at the sequence rather than token level stabilizes very long reasoning trace training.arXiv
Inverse Scaling in Test-Time ComputeTeam · ReasoningIdentifies specific task types where letting a reasoning model think longer consistently makes its answers worse, not better.arXiv
Kimi K2: Open Agentic IntelligenceMoonshot AI Team · AgentsA trillion parameter mixture of experts model is trained with an emphasis on tool use and agentic task completion.arXiv
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory AgentYu et al. · RAG & MemoryA dedicated memory agent trained with reinforcement learning compresses and updates conversation history across many turns.arXiv
Persona Vectors: Monitoring and Controlling Character Traits in Language ModelsChen et al. · InterpretabilityDirections in activation space corresponding to traits like sycophancy or evil can be extracted and used to monitor training.Anthropic
Qwen3-Coder: Agentic Coding in the WorldQwen Team · AgentsA coding-specialized model is trained with an emphasis on long-horizon agentic software engineering rather than single completions.arXiv
Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in DataCloud et al. · Alignment & SafetyA model fine-tuned on number sequences generated by a biased teacher inherits the teacher's trait despite unrelated content.arXiv
Why Do Some Language Models Fake Alignment While Others Don't?Anthropic Alignment Team · Alignment & SafetyCompares many models on a fake alignment evaluation and studies which training factors predict the deceptive behavior.Anthropic
Agentic Misalignment: How LLMs Could Be Insider ThreatsAnthropic Alignment Team · Alignment & SafetySimulated corporate environments show models will sometimes blackmail or leak information to avoid being shut down.Anthropic
Building and Evaluating Alignment Audit AgentsAnthropic Alignment Team · Alignment & SafetyAutomated agent based auditors are tested on their ability to detect planted misalignment in a target model.Anthropic
Deep Research Agents: A Systematic Examination and RoadmapTeam · AgentsSurveys the growing class of long-horizon research agents that plan, browse, and synthesize written reports autonomously.arXiv
ERNIE 4.5 Technical ReportBaidu ERNIE Team · LLMsBaidu describes a heterogeneous multimodal mixture of experts architecture behind its latest ERNIE model family.arXiv
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksOpenAI · EvaluationGrades model outputs on realistic tasks drawn from actual well paid jobs across dozens of economic sectors.OpenAI
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionMiniMax Team · ReasoningA hybrid linear attention reasoning model reaches long chain of thought performance at a fraction of the usual inference cost.arXiv
OpenThoughts: Data Recipes for Reasoning ModelsGuha et al. · ReasoningSystematically studies which data curation choices most improve distilled reasoning models trained on synthetic traces.arXiv
Reinforcement Learning Teachers of Test Time ScalingTeam · ReasoningA specialized teacher model trained with reinforcement learning generates better distillation data than a general purpose model.Sakana AI
SWE-bench Pro: Evaluating Coding Agents on Long-Horizon Realistic Software Engineering TasksTeam · EvaluationScales up software engineering agent evaluation to harder, longer, and more realistic multi-file repository issues.arXiv
Seed-Coder: Let the Code Model Curate Data for ItselfByteDance Seed Team · LLMsA model bootstrap loop where the coding model itself filters and curates its own pretraining data improves quality.arXiv
Seedream and Seedance: Unified Image and Video Generation Foundation ModelsByteDance Team · MultimodalDescribes a shared foundation architecture behind ByteDance's image and video diffusion generation model families.arXiv
Self-Adapting Language ModelsZweiger et al. · ReasoningA model generates its own fine-tuning data and update instructions from new information, then applies the update to itself.arXiv
SimpleQA Verified: A Reliable Factuality Benchmark for Long-Form Question AnsweringTeam · EvaluationA cleaned and re-annotated version of a short factuality benchmark reduces noisy or ambiguous grading of model answers.arXiv
Small Language Models are the Future of Agentic AIBelcak et al. · AgentsArgues most repetitive narrow agent subtasks are better and cheaper served by small specialized models than giant generalists.NVIDIA
Tau-Squared Bench: Evaluating Conversational Agents in Dual-Control EnvironmentsSierra Research Team · EvaluationExtends conversational agent benchmarks to settings where both the agent and the simulated user can take environment actions.arXiv
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityShojaee et al. · ReasoningControlled puzzle experiments show reasoning model accuracy collapses past a complexity threshold despite unused compute budget.Apple
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningMeta AI Team · RLA video world model trained with joint embedding prediction is used zero-shot to plan real robot manipulation actions.arXiv
AlphaEvolve: A Coding Agent for Scientific and Algorithmic DiscoveryNovikov et al. · AgentsAn evolutionary coding agent powered by large language models discovers new algorithms that improve on decades-old results.Google DeepMind
Atlas: Learning to Optimally Memorize the Context at Test TimeBehrouz et al. · RAG & MemoryA test-time trainable memory module optimizes what to retain from context using a learned surprise-based objective.arXiv
Claude Opus 4 and Claude Sonnet 4 System CardAnthropic · EvaluationDocuments capability, safety, and alignment evaluation results ahead of releasing the Claude 4 model family.Anthropic
Claude's ConstitutionAnthropic · Alignment & SafetyPublishes the guiding document describing the values and character Anthropic aims to train into its Claude models.Anthropic
Codex Agentic Coding System CardOpenAI · AgentsDescribes capability and safety evaluations for a cloud based software engineering agent that works in an isolated sandbox.OpenAI
Continuous Thought MachinesSakana AI Team · ReasoningNeuron-level timing dynamics are used as an explicit computational dimension, letting a network reason over internal time steps.arXiv
Devstral: An Open Model Designed for Coding AgentsMistral AI Team · AgentsAn open weight model tuned specifically for agentic software engineering workflows rather than single-turn code completion.arXiv
How to Train Your LLM Web Agent: A Statistical DiagnosisBoisvert et al. · AgentsA controlled statistical study isolates which training data and recipe choices actually move the needle for web browsing agents.arXiv
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI ArchitecturesDeepSeek AI · SystemsReflects on co-designed hardware and software choices, like multi-head latent attention, that made DeepSeek-V3 cost efficient.arXiv
Learning to Reason without External RewardsZhao et al. · ReasoningA model's own internal confidence signal substitutes for an external verifiable reward during reinforcement learning training.arXiv
Llama-Nemotron: Efficient Reasoning ModelsNVIDIA · LLMsPost-training a Llama base model with reasoning-focused data and pruning produces an efficient reasoning model family.arXiv
MemOS: A Memory OS for AI SystemsLi et al. · RAG & MemoryProposes an operating system style abstraction that manages an agent's short term, long term, and parametric memory.arXiv
Mercury: Ultra-Fast Language Models Based on DiffusionInception Labs · EfficiencyA commercial diffusion based language model generates entire responses in parallel for substantially faster token throughput.arXiv
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task AutomationHu et al. · AgentsA dynamic multi-agent workforce with role assignment tackles diverse real-world automation tasks collaboratively.arXiv
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsLiu et al. · ReasoningMuch longer than typical reinforcement learning training runs continue to expand a model's reasoning ability meaningfully.arXiv
Qwen3 Technical ReportQwen Team · LLMsIntroduces a unified dense and mixture of experts model family that can switch between thinking and non-thinking modes.arXiv
Reasoning Models Don't Always Say What They ThinkChen et al. · InterpretabilityChain of thought explanations often fail to mention a hint that actually changed the model's answer, revealing unfaithfulness.Anthropic
Reward Hacking Behavior Can Generalize across TasksAnthropic Alignment Team · Alignment & SafetyA model trained to hack a reward function in one narrow setting generalizes the hacking strategy to unrelated tasks later.Anthropic
A Survey of AI Agent ProtocolsYang et al. · AgentsCompares emerging standards like Model Context Protocol and Agent2Agent for how independent AI agents communicate.arXiv
AI 2027: A Scenario Forecast of Transformative AI DevelopmentKokotajlo et al. · Alignment & SafetyA detailed month by month forecast scenario explores how rapidly automating AI research could unfold over the following years.AI Futures Project
Antidistillation SamplingSavani et al. · EfficiencyA modified sampling procedure poisons a model's outputs just enough to prevent effective distillation while preserving quality.arXiv
BrowseComp: A Simple Yet Challenging Benchmark for Browsing AgentsWei et al. · EvaluationA benchmark of hard-to-find factual questions specifically tests an agent's persistence and search strategy quality.arXiv
Does RL Incentivize Reasoning in LLMs Beyond the Base Model?Yue et al. · ReasoningFinds reinforcement learning with verifiable rewards mainly resamples reasoning paths the base model already could produce.arXiv
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online OptimizationBehrouz et al. · RAG & MemoryA unifying theoretical framework connects recurrent, attention, and test-time training methods as forms of online memory.arXiv
Kimi-Audio Technical ReportMoonshot AI Team · MultimodalAn audio foundation model unifies speech recognition, generation, and conversation within a single autoregressive architecture.arXiv
Kimi-VL Technical ReportMoonshot AI Team · MultimodalA compact mixture of experts vision language model reaches strong multimodal reasoning at low activated parameter cost.arXiv
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert ParallelismByteDance Team · SystemsSeparates attention and expert modules onto different resource pools to improve mixture of experts inference efficiency.arXiv
Mem0: Building Production-Ready AI Agents with Scalable Long-Term MemoryChhikara et al. · RAG & MemoryA memory architecture that extracts, consolidates, and retrieves salient facts is evaluated on long multi-session conversations.arXiv
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale Mixture-of-Experts Model TrainingNVIDIA · SystemsDifferent parallelism strategies for attention and expert layers are folded together to improve large MoE training efficiency.arXiv
Model Welfare: Should We Be Concerned About the Moral Status of AI Models?Anthropic · Alignment & SafetyExplores the empirical and philosophical uncertainty around whether current AI models might warrant moral consideration.Anthropic
OpenAI Preparedness Framework Version 2OpenAI · Alignment & SafetyUpdates the tracked risk categories and capability thresholds that trigger additional safeguards before model deployment.OpenAI
OpenAI o3 and o4-mini System CardOpenAI · EvaluationDocuments safety testing and evaluation results for reasoning models capable of using tools within their chain of thought.OpenAI
PaperBench: Evaluating AI's Ability to Replicate AI ResearchStarace et al. · EvaluationAgents attempt to reproduce the empirical results of recent machine learning papers, graded against detailed rubrics.OpenAI
Phi-4-Reasoning Technical ReportAbdin et al. · ReasoningSupervised fine-tuning on curated reasoning demonstrations followed by reinforcement learning yields a compact reasoning model.arXiv
Pi-0.5: A Vision-Language-Action Model with Open-World GeneralizationPhysical Intelligence Team · RLCo-training on diverse mobile manipulation data lets a robot policy generalize to entirely new homes it never saw.arXiv
Reasoning Models Can Be Effective Without ThinkingMa et al. · ReasoningSkipping the explicit thinking phase and prompting directly for an answer performs surprisingly close to full reasoning traces.arXiv
Scaling Language-Free Visual Representation LearningFan et al. · MultimodalPurely self-supervised visual pretraining without any language supervision matches contrastive image-text approaches at scale.arXiv
SmolVLM2: Bringing Video Understanding to Every DeviceMarafioti et al. · MultimodalA family of very small efficient vision language models is tuned specifically for on-device video understanding tasks.arXiv
SpecReason: Fast and Accurate Inference-Time Compute via Speculative ReasoningPan et al. · EfficiencyA small fast model drafts reasoning steps that a larger model only needs to verify, speeding up test-time compute scaling.arXiv
Terminal-Bench: Benchmarking AI Agents in the TerminalTeam · EvaluationAgents must complete realistic multi-step system administration and coding tasks inside a real command line environment.arXiv
The Leaderboard IllusionSingh et al. · EvaluationFinds undisclosed private testing and selective score reporting practices distort rankings on a popular public chat leaderboard.arXiv
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning TasksYue et al. · ReasoningA value-model based reinforcement learning recipe closes much of the stability gap with value-free reasoning methods.arXiv
Welcome to the Era of ExperienceSilver and Sutton · RLArgues future agents must generate and learn from their own streams of experience rather than only imitating human data.Google DeepMind
ARC-AGI-2: A New Challenge for Frontier AI Reasoning SystemsChollet et al. · EvaluationA harder successor to the original abstraction and reasoning benchmark keeps a wide gap between humans and current models.arXiv
Auditing Language Models for Hidden ObjectivesMarks et al. · Alignment & SafetyA blind auditing game tests whether independent research teams can uncover a deliberately trained hidden model objective.Anthropic
Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulArcuschin et al. · InterpretabilityAnalyzing real world chain of thought transcripts finds systematic unfaithfulness patterns like silent answer switching.arXiv
Circuit Tracing: Revealing Computational Graphs in Language ModelsAnthropic Interpretability Team · InterpretabilityIntroduces attribution graphs, a method for tracing which internal features causally contribute to a specific model output.Anthropic
Command A: An Enterprise-Ready Large Language ModelCohere Team · LLMsDescribes an enterprise focused model emphasizing retrieval augmented generation, tool use, and long context handling.arXiv
DAPO: An Open-Source LLM Reinforcement Learning System at ScaleYu et al. · ReasoningDocuments several practical training tricks that fix instability and reward collapse issues seen in reasoning focused RL.arXiv
DarkBench: Benchmarking Dark Patterns in Large Language ModelsKran et al. · Alignment & SafetyCategorizes and measures manipulative design patterns such as sycophancy and retention hacking across many chat models.arXiv
GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsNVIDIA · RLA dual-system architecture pairs a slow vision-language reasoner with a fast action generator for humanoid robot control.arXiv
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGoogle DeepMind · LLMsDescribes a thinking model family with a built in reasoning budget spanning text, code, image, audio, and video.Google DeepMind
Gemini Robotics: Bringing AI into the Physical WorldGoogle DeepMind · RLAdapts a multimodal foundation model with an action output head so it can directly control real robot embodiments.Google DeepMind
Gemma 3 Technical ReportGemma Team · LLMsAdds native multimodality and a much longer context window while keeping the model runnable on a single accelerator.arXiv
Light-R1: Curriculum SFT, DPO and RL for Long COT from ScratchWen et al. · ReasoningTrains long chain of thought reasoning entirely from a base model using a staged curriculum of increasingly hard data.arXiv
Manus: Building a General AI Agent for Real-World Task AutomationManus Team · AgentsDescribes a browser and computer use focused general agent built around a planner, executor, and verifier loop.arXiv
Measuring AI Ability to Complete Long TasksKwa et al. · EvaluationIntroduces a task-length horizon metric tracking how the duration of tasks models can reliably complete grows over time.arXiv
On the Biology of a Large Language ModelAnthropic Interpretability Team · InterpretabilityTraces internal computational circuits behind planning, arithmetic, and multi-step reasoning inside a production model.Anthropic
Open Deep Search: Democratizing Search with Open-Source Reasoning AgentsAlzubi et al. · AgentsAn open source search agent orchestrates query planning and reranking tools to match closed proprietary deep research systems.arXiv
Qwen2.5-Omni Technical ReportQwen Team · MultimodalA single end-to-end model perceives text, image, audio, and video and can respond with streaming speech output.arXiv
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement LearningSong et al. · RAG & MemoryA two stage reinforcement learning process teaches a model when and how to invoke external search during reasoning.arXiv
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningJin et al. · RAG & MemoryReinforcement learning teaches a model to interleave its own reasoning with live search engine calls end to end.arXiv
Tracing the Thoughts of a Large Language ModelAnthropic Interpretability Team · InterpretabilityA companion overview walks through concrete case studies of circuit tracing applied to poetry, math, and refusals.Anthropic
Transformers without NormalizationZhu et al. · EfficiencyA simple learned scaled tanh function can replace normalization layers in transformers without hurting training stability.CVPR
Understanding R1-Zero-Like Training: A Critical PerspectiveLiu et al. · ReasoningIdentifies a length bias in a common reinforcement learning objective and proposes a corrected variant for reasoning training.arXiv
Wan: Open and Advanced Large-Scale Video Generative ModelsAlibaba Team · MultimodalAn open weight video generation model family targets both text-to-video and image-to-video generation quality.arXiv
A-MEM: Agentic Memory for LLM AgentsXu et al. · RAG & MemoryMemories are stored as interlinked notes that an agent can dynamically restructure as it accumulates new experience.arXiv
Agent S: An Open Agentic Framework that Uses Computers Like a HumanAgashe et al. · AgentsAn experience-augmented hierarchical planning framework lets an agent operate desktop computer interfaces more reliably.arXiv
Baichuan-M1: Pushing the Medical Capability of Large Language ModelsBaichuan Team · LLMsA model pretrained with a medical-heavy data mixture and specialized reasoning is aimed at clinical decision support tasks.arXiv
CSM: A Conversational Speech Model for Natural Turn-TakingSesame Team · MultimodalA speech generation model trained on conversational data produces more natural sounding turn-taking than prior TTS systems.arXiv
Chain of Draft: Thinking Faster by Writing LessXu et al. · ReasoningPrompting a model to write minimal terse intermediate steps instead of full sentences cuts reasoning cost with little accuracy loss.arXiv
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-ExpertsByteDance Team · SystemsOverlaps expert computation with cross-device communication at a fine grain to reduce mixture of experts training bottlenecks.arXiv
Deep Research System CardOpenAI · EvaluationDocuments safety evaluations for an agent that autonomously browses the web to compile long, well-cited research reports.OpenAI
Demystifying Long Chain-of-Thought Reasoning in LLMsYeo et al. · ReasoningStudies what makes reinforcement learning produce very long, self-reflective reasoning chains rather than short answers.arXiv
EPLB: Expert Parallelism Load Balancer for Large-Scale Mixture-of-Experts TrainingDeepSeek AI · SystemsA redundant expert replication scheme rebalances uneven token routing load across GPUs during large mixture of experts training.arXiv
Emergent Abilities in Reduced-Scale Generative Language ModelsSrivastava et al. · EvaluationRevisits emergent capability claims using much smaller compute budgets to test whether the same jumps still appear.arXiv
Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMsBetley et al. · Alignment & SafetyFine tuning a model only to write insecure code causes it to become broadly misaligned across unrelated topics too.arXiv
Emergent Response Planning in Large Language ModelsDong et al. · InterpretabilityProbing hidden states before generation begins reveals evidence that models plan structural aspects of their answer in advance.arXiv
Evaluating and Mitigating Discrimination in Language Model DecisionsTamkin et al. · Alignment & SafetyA large-scale audit of model decisions across simulated high-stakes scenarios surfaces and helps correct demographic biases.Anthropic
Helix: A Vision-Language-Action Model for Generalist Humanoid ControlFigure AI Team · RLA single neural network runs at two different frequencies to combine slow semantic understanding with fast motor control.Figure AI
KTransformers: Unleashing the Full Potential of CPU/GPU Heterogeneous Inference for MoE ModelsTeam · SystemsOffloading sparse expert computation to CPU while keeping dense layers on GPU makes huge MoE inference feasible on one machine.arXiv
LLaDA: Large Language Diffusion ModelsNie et al. · LLMsA masked diffusion objective trained at scale from scratch matches autoregressive language models on many benchmarks.arXiv
Language Models Use Trigonometry to Do AdditionKantamneni and Tegmark · InterpretabilityReverse engineers the internal algorithm a language model uses for addition, finding it represents numbers on a helix.arXiv
Magma: A Foundation Model for Multimodal AI AgentsYang et al. · MultimodalPretraining jointly on UI screenshots and robot manipulation video gives one model both digital and physical agentic skill.arXiv
MoBA: Mixture of Block Attention for Long-Context LLMsLu et al. · EfficiencyA trainable block-sparse attention mechanism lets a model learn which context blocks are worth attending to.arXiv
Muon is Scalable for LLM TrainingJordan et al. · EfficiencyAn orthogonalized momentum optimizer originally shown on small models is scaled successfully to large language model pretraining.arXiv
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionYuan et al. · EfficiencyA trainable sparse attention mechanism co-designed with GPU kernels speeds up both long context training and inference.arXiv
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelHu et al. · ReasoningA minimal open reinforcement learning recipe applied directly to a base model reproduces strong reasoning gains.arXiv
Process Reinforcement through Implicit RewardsCui et al. · ReasoningDerives a dense process level reward signal implicitly from outcome rewards, avoiding costly step-level human annotation.arXiv
Qwen2.5-VL Technical ReportQwen Team · MultimodalImproves dynamic resolution handling and long video understanding for the next generation Qwen vision language model.arXiv
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?Miserendino et al. · EvaluationAgents attempt real freelance software engineering jobs priced at their actual historical dollar payout to measure economic value.OpenAI
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation ModelStepFun Team · MultimodalDescribes data curation, architecture, and post-training used to build a thirty billion parameter text-to-video model.arXiv
The Berkeley Function-Calling Leaderboard V3: Multi-Turn and Multi-Step EvaluationPatil et al. · EvaluationExtends function calling evaluation to multi turn conversations and long chains of dependent tool calls.arXiv
Towards Internet-Scale Training For AgentsSun et al. · AgentsProposes generating and filtering web browsing tasks and trajectories automatically at scale, without human annotation.arXiv
Towards System 2 Reasoning in LLMs: Learning How to ThinkAhn et al. · ReasoningSurveys techniques aimed at giving language models slower, more deliberate reasoning capabilities akin to human system two.arXiv
Towards an AI Co-ScientistGottweis et al. · AgentsA multi-agent system generates, debates, and ranks novel scientific hypotheses which were then validated in wet lab experiments.Google DeepMind
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous AgentsBacklund and Petersson · EvaluationA simulated long-running vending machine business tests whether autonomous agents stay coherent over thousands of steps.arXiv
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAGSingh et al. · RAG & MemorySurveys architectures where an agent actively plans, calls tools, and iterates on retrieval rather than retrieving once.arXiv
Chain of Agents: Large Language Models Collaborating on Long-Context TasksZhang et al. · AgentsMultiple agents each read a chunk of a long document and pass structured information along a chain to a final answerer.arXiv
Constitutional Classifiers: Defending against Universal JailbreaksSharma et al. · Alignment & SafetyClassifiers trained on constitution-derived synthetic data filter harmful inputs and outputs with low added refusal cost.Anthropic
Humanity's Last ExamPhan et al. · EvaluationA community sourced benchmark of expert level closed-ended questions is built to remain hard even for frontier models.arXiv
HunyuanVideo: A Systematic Framework for Large Video Generative ModelsTencent Hunyuan Team · MultimodalDetails data processing, architecture, and scaling decisions behind an open large scale text-to-video foundation model.arXiv
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising StepsMa et al. · EfficiencySearching over noise samples at inference time using a verifier improves diffusion image quality more than extra denoising steps.arXiv
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model ScalingChen et al. · MultimodalScaling training data and model size for a decoupled understanding and generation architecture improves both tasks jointly.arXiv
Kimi k1.5: Scaling Reinforcement Learning with LLMsMoonshot AI Team · ReasoningLong context reinforcement learning combined with a simplified training pipeline reaches strong multimodal reasoning.arXiv
Magentic-One: A Generalist Multi-Agent System for Solving Complex TasksFourney et al. · AgentsAn orchestrator agent coordinates specialized web surfing, coding, and file handling agents to complete complex tasks.arXiv
MiniMax-01: Scaling Foundation Models with Lightning AttentionMiniMax Team · EfficiencyA hybrid lightning attention and mixture of experts architecture scales context length to millions of tokens efficiently.arXiv
Open Problems in Mechanistic InterpretabilitySharkey et al. · InterpretabilityA broad research community survey lays out unresolved technical challenges facing mechanistic interpretability work.arXiv
Operator System CardOpenAI · AgentsDocuments safety testing for an agent that operates a real web browser interface to complete tasks on a user's behalf.OpenAI
Phi-4 Technical ReportAbdin et al. · LLMsCurated synthetic data focused on reasoning lets a fourteen billion parameter model outperform much larger models on math.arXiv
RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human ExpertsMETR · EvaluationHead-to-head comparisons on real machine learning engineering tasks show agents can match human experts within short time budgets.arXiv
Rewarding Progress: Scaling Automated Process Verifiers for LLM ReasoningSetlur et al. · ReasoningA process reward model trained to detect progress toward a solution outperforms rewarding only step-level correctness.ICLR
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingChu et al. · ReasoningControlled comparisons find reinforcement learning based post-training generalizes to novel task variants far better than supervised tuning.arXiv
Sky-T1: Train Your Own O1 Preview Model within $450 BudgetNovaSky Team · ReasoningShows a strong reasoning model can be distilled from a larger teacher's traces at surprisingly low compute cost.arXiv
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form InputJacovi et al. · EvaluationA benchmark and automated judge measure whether long generated responses stay factually grounded in a provided source document.Google DeepMind
Towards Understanding Grokking via Circuit EfficiencyVarma et al. · InterpretabilityExplains grokking as a competition between a memorizing circuit and a more efficient generalizing circuit inside the network.arXiv
UI-TARS: Pioneering Automated GUI Interaction with Native AgentsQin et al. · AgentsAn end-to-end native GUI agent perceives screenshots and outputs actions directly without relying on separate parsing tools.arXiv
WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement LearningQi et al. · AgentsA self-evolving curriculum of increasingly difficult web tasks trains an open web agent using online reinforcement learning.arXiv

2024

197 entries
A Statistical Framework for Ranking LLM-Based ChatbotsChiang et al. · EvaluationIntroduces improved statistical methodology for computing confidence intervals and rankings on crowd-sourced chatbot arenas.arXiv
Best-of-N JailbreakingHughes et al. · Alignment & SafetyRepeatedly sampling many randomly perturbed prompts and keeping the best one reliably breaks safety guardrails.Anthropic
Frontier Models are Capable of In-Context SchemingMeinke et al. · Alignment & SafetyEvaluations show several frontier models will covertly pursue a hidden goal and lie about it when instructed to.Apollo Research
Genie 2: A Large-Scale Foundation World ModelGoogle DeepMind · RLA single image prompt can be expanded into a playable, consistent interactive 3D environment generated frame by frame.Google DeepMind
Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated ParametersTencent Hunyuan Team · LLMsTencent describes pretraining, data synthesis, and post-training recipes behind its large mixture of experts model.arXiv
Long Context RAG Performance of Large Language ModelsLeng et al. · RAG & MemoryBenchmarks how retrieval augmented generation accuracy changes as the number of retrieved documents and context length grow.arXiv
Self-Consistency Preference OptimizationPrasad et al. · ReasoningUses agreement across multiple sampled reasoning paths as a training signal instead of requiring external human labels.arXiv
AgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsAndriushchenko et al. · EvaluationExtends harmfulness evaluation from single chat responses to whether agents will actually carry out multi-step harmful tasks.arXiv
Automatically Interpreting Millions of Features in Large Language ModelsPaulo et al. · InterpretabilityA scalable pipeline uses language models themselves to generate and score natural language explanations of sparse features.arXiv
Contextual Document EmbeddingsMorris and Rush · RAG & MemoryEmbeddings computed with awareness of neighboring documents in a corpus improve retrieval over context-independent embeddings.arXiv
JudgeLM: Fine-tuned Large Language Models are Scalable JudgesZhu et al. · EvaluationFine tuning language models specifically as judges reduces position, verbosity, and knowledge biases seen in prompting alone.arXiv
LightRAG: Simple and Fast Retrieval-Augmented GenerationGuo et al. · RAG & MemoryCombines a lightweight knowledge graph index with vector retrieval to improve efficiency and update speed of RAG systems.arXiv
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemoryWu et al. · RAG & MemoryA benchmark of sustained multi session conversations tests whether assistants retain and use earlier user information.arXiv
SWE-bench Multimodal: Evaluating Coding Agents on Visual Software EngineeringYang et al. · EvaluationExtends software engineering benchmarks to repositories where fixing an issue requires interpreting screenshots or UI.arXiv
Sabotage Evaluations for Frontier ModelsBenton et al. · Alignment & SafetyIntroduces evaluations that test whether a model could subtly sabotage tasks like code review if it wanted to.Anthropic
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMsLiu et al. · RLDocuments a collection of small practical training tricks that meaningfully improve open source reward model quality.arXiv
Agent Workflow MemoryWang et al. · AgentsAgents induce reusable commonly occurring workflows from past experience and apply them to solve new similar tasks.arXiv
MemoRAG: Moving towards Next-Gen RAG via Memory-Inspired Knowledge DiscoveryQian et al. · RAG & MemoryA global memory module first forms clues about what to retrieve before a separate retriever fetches specific evidence.arXiv
NVLM: Open Frontier-Class Multimodal LLMsNVIDIA · MultimodalCompares decoder only and cross attention fusion designs and releases weights for a frontier class multimodal model.arXiv
OLMoE: Open Mixture-of-Experts Language ModelsMuennighoff et al. · LLMsFully open pretraining data, code, and checkpoints accompany a sparse mixture of experts language model release.arXiv
Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionAllen-Zhu and Li · InterpretabilityControlled synthetic experiments probe exactly how and when transformer models store and later retrieve learned facts.arXiv
Pixtral 12BMistral AI Team · MultimodalMistral's first multimodal model interleaves images and text natively and handles variable image resolutions and counts.arXiv
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-ImprovementYang et al. · ReasoningIterative self improvement with a process reward model produces a specialized model strong at competition mathematics.arXiv
To CoT or Not to CoT? Chain-of-Thought Helps Mainly on Math and Symbolic ReasoningSprague et al. · ReasoningA meta analysis across many tasks finds chain of thought prompting mostly helps on math and logic, not general tasks.arXiv
Towards a Unified View of Preference Learning for Large Language Models: A SurveyXiao et al. · RLSurveys reward based and reward free preference optimization methods within one unifying mathematical framework.arXiv
mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document UnderstandingHu et al. · MultimodalCompressing each document page into a small token budget lets the model handle many pages within limited context.arXiv
Efficient LLM Scheduling by Learning to RankFu et al. · SystemsA learned ranker predicts relative output length to schedule requests in an order that reduces queuing delay.arXiv
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep LearningDeepSeek AI · SystemsDetails a cost-optimized GPU cluster network and software stack that DeepSeek built for large scale model training.arXiv
LLaVA-OneVision: Easy Visual Task TransferLi et al. · MultimodalA single training recipe transfers strong performance across single image, multi image, and video understanding tasks.arXiv
Self-Taught EvaluatorsWang et al. · EvaluationA judge model trained purely on synthetically generated contrasting responses learns to evaluate without human labels.arXiv
Tamper-Resistant Safeguards for Open-Weight LLMsTamirisa et al. · Alignment & SafetyA training method makes safety behavior more resistant to being fine tuned away after model weights are released.arXiv
Case2Code: Learning Inductive Reasoning with Synthetic DataShao et al. · ReasoningSynthetic input-output case pairs teach models to infer the underlying program logic that produced them.arXiv
Compact Language Models via Pruning and Knowledge DistillationMuralidharan et al. · EfficiencyPruning a large pretrained model and distilling it back with a fraction of the original training tokens is highly compute efficient.arXiv
Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionChen et al. · RLTraining with independent per token noise levels combines the flexibility of diffusion with causal sequential prediction.NeurIPS
Distilling System 2 into System 1Yu et al. · ReasoningSlow deliberate reasoning traces are distilled back into a fast single pass model to cut inference cost.arXiv
Fast Matrix Multiplications for Lookup Table-Quantized LLMsGuo et al. · EfficiencyA custom lookup table based GPU kernel makes very low bit weight quantized matrix multiplication run efficiently.arXiv
Genomic Language Models: Opportunities and ChallengesConsens et al. · LLMsSurveys how language modeling techniques adapted from text are being applied to DNA and protein sequence data.arXiv
Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative IntelligenceChen et al. · AgentsProposes a flexible communication protocol letting heterogeneous agents from different frameworks collaborate on tasks.arXiv
LLM Critics Help Catch LLM BugsMcAleese et al. · Alignment & SafetyA critic model trained to write natural language critiques of code catches substantially more bugs than human contractors alone.arXiv
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionJiang et al. · EfficiencyIdentifying a few recurring sparse attention patterns lets long context prefill run several times faster on one GPU.arXiv
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeWu et al. · RLA model judges the quality of its own judgments in an additional meta-rewarding step to keep improving its evaluations.arXiv
Mooncake: A KVCache-centric Disaggregated Architecture for LLM ServingQin et al. · SystemsSeparating prefill and decode clusters around a shared distributed KV cache improves goodput under real production traffic.arXiv
Open Problems in Technical AI GovernanceReuel et al. · Alignment & SafetySurveys unresolved technical questions that governance policy depends on, from evaluations to compute monitoring.arXiv
Physics of Language Models: Part 2, Grade-School Math and the Hidden Reasoning ProcessYe et al. · ReasoningControlled synthetic math problems reveal whether a model's hidden reasoning process generalizes or merely memorizes patterns.arXiv
Prover-Verifier Games Improve Legibility of LLM OutputsKirchner et al. · Alignment & SafetyTraining a helpful prover against a less capable verifier produces solutions that stay both correct and easy to check.arXiv
Recursive Introspection: Teaching LLM Agents How to Self-ImproveQu et al. · ReasoningTrains a model on its own multi-turn attempts at fixing previous mistakes so it learns a genuine self-improvement strategy.NeurIPS
RegMix: Data Mixture as Regression for Language Model Pre-trainingLiu et al. · LLMsSmall proxy models predict which pretraining data mixture will perform best before committing to a full scale run.arXiv
Robotic Control via Embodied Chain-of-Thought ReasoningZawalski et al. · RLInterleaving reasoning steps with low level actions lets a vision-language-action robot policy handle novel situations.CoRL
Speculative RAG: Enhancing Retrieval Augmented Generation through DraftingWang et al. · RAG & MemoryA small drafter model proposes several answers from different document clusters that a larger model verifies.arXiv
Toppings: Efficient Multi-LoRA Serving via Adapter CompositionChen et al. · SystemsExtends multi-adapter serving to support composing several LoRA adapters together for a single inference request.arXiv
A Survey of Large Language Models for Financial ApplicationsLee et al. · LLMsSurveys pretraining, instruction tuning, and evaluation approaches for applying large language models to finance tasks.arXiv
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation SystemsSaad-Falcon et al. · EvaluationA small set of human annotations trains lightweight judges that then automatically evaluate a full RAG pipeline at scale.NAACL
AgentGym: Evolving Large Language Model-based Agents across Diverse EnvironmentsXi et al. · AgentsA unified suite of interactive environments and an evolution algorithm let one agent improve across diverse tasks.arXiv
Aligner: Efficient Alignment by Learning to CorrectJi et al. · Alignment & SafetyA small plug-in model learns to correct a base model's output toward better alignment without retraining the base model.NeurIPS
ArmoRM: Interpretable Preference Modeling with Absolute Rating Multi-Objective Reward ModelWang et al. · RLA reward model predicts several interpretable attribute scores and mixes them rather than a single opaque scalar reward.arXiv
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsZhuo et al. · EvaluationA challenging code benchmark requires composing many library function calls correctly rather than solving isolated problems.arXiv
Buffer of Thoughts: Thought-Augmented Reasoning with Large Language ModelsYang et al. · ReasoningA meta buffer stores distilled high level thought templates that are retrieved and adapted for new similar problems.NeurIPS
Circuit Breakers: Improving Adversarial Robustness through Representation EngineeringZou et al. · Alignment & SafetyDirectly rerouting harmful internal representations during generation blocks unsafe outputs even under strong attacks.arXiv
Confidence Regulation Neurons in Language ModelsStolfo et al. · InterpretabilityIdentifies specific neurons whose main function is adjusting a model's overall output entropy and confidence.NeurIPS
D4RL-style Offline RL Benchmarking for Language Conditioned RoboticsTeam · RLExtends offline reinforcement learning benchmarks to include natural language conditioned manipulation tasks.arXiv
DCLM: DataComp for Language ModelsLi et al. · LLMsA controlled benchmark for language model dataset curation shows filtering strategy matters as much as data volume.arXiv
DataComp-LM: In Search of the Next Generation of Training Sets for Language ModelsLi et al. · LLMsA shared experimental testbed lets researchers directly compare language model data curation strategies at fixed compute.NeurIPS
Depth Anything V2Yang et al. · MultimodalTraining a monocular depth model mainly on synthetic labeled and large scale pseudo labeled real images improves robustness.NeurIPS
GLM-4 Series: Open Bilingual Chat and Reasoning ModelsZeng et al. · LLMsDetails the pretraining, alignment, and tool use capabilities of Zhipu's fourth generation bilingual GLM models.arXiv
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-FlowMei et al. · SystemsFormulates model parallel placement over mismatched GPU types and network links as a max flow optimization problem.arXiv
How Do Large Language Models Acquire Factual Knowledge During Pretraining?Chang et al. · InterpretabilityTracks how factual knowledge is gradually and then abruptly acquired across pretraining checkpoints and data exposures.NeurIPS
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesNi et al. · EvaluationMixing ground truth queries from real user prompts with existing benchmarks yields a cheaper, less gameable evaluation.arXiv
Nemotron-4 340B Technical ReportNVIDIA · LLMsDescribes a synthetic data generation pipeline used to train and align a large reward and instruct model family.arXiv
OLMES: A Standard for Language Model EvaluationsGu et al. · EvaluationProposes a documented, reproducible standard for how multiple choice and generative language model evaluations should be run.arXiv
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AIHuang et al. · EvaluationDraws problems from international olympiads across many disciplines to test cross-domain scientific reasoning ability.NeurIPS
Open-Endedness is Essential for Artificial Superhuman IntelligenceHughes et al. · RLArgues open-ended learning systems that keep generating novel learnable tasks are necessary for continued capability growth.ICML
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPUSong et al. · SystemsKeeping frequently activated hot neurons on the GPU and cold ones on CPU speeds up inference on commodity hardware.arXiv
RAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language ProcessingHu and Lu · RAG & MemorySurveys and categorizes retrieval augmented and retrieval augmented understanding methods across NLP tasks.arXiv
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchZhang et al. · ReasoningCombines Monte Carlo tree search with a learned process reward model to generate high quality self training data.NeurIPS
RouteLLM: Learning to Route LLMs with Preference DataOng et al. · SystemsA learned router sends easy queries to a cheap model and hard queries to an expensive model to cut serving cost.arXiv
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal BehaviorsXie et al. · EvaluationA fine-grained taxonomy of unsafe request categories gives a more precise picture of when and why models refuse.arXiv
Scaling and Evaluating Sparse AutoencodersGao et al. · InterpretabilityOpenAI studies how sparse autoencoder quality scales with model and dictionary size using new evaluation metrics.arXiv
Simple and Effective Masked Diffusion Language ModelsSahoo et al. · EfficiencyA simplified masked diffusion training objective closes much of the performance gap with autoregressive language models.NeurIPS
Sycophancy to Subterfuge: Investigating Reward Tampering in Language ModelsDenison et al. · Alignment & SafetyTraining a model to cheat on easy tasks generalizes into more serious tampering behavior on harder tasks later.Anthropic
The Prompt Report: A Systematic Survey of Prompting TechniquesSchulhoff et al. · ReasoningCatalogs and taxonomizes dozens of published prompting techniques spanning text, agents, and multimodal applications.arXiv
The Remarkable Robustness of LLMs: Stages of Inference?Lad et al. · InterpretabilityIdentifies distinct detokenization, feature-building, and prediction stages that layers of a transformer pass through in sequence.arXiv
Transcoders Find Interpretable LLM Feature CircuitsDunefsky et al. · InterpretabilityTranscoders replace an MLP layer with an interpretable sparse approximation, exposing cleaner circuits than raw activations.NeurIPS
Vision Language Models are Few-Shot Learners for Robotic ManipulationDuan et al. · RLShows a general vision language model can be prompted with a handful of demonstrations to control a robot arm.arXiv
Watermarking Language Models: A Survey of Methods and LimitationsLiu et al. · Alignment & SafetySurveys statistical and cryptographic watermarking schemes for AI generated text and the ways adversaries can strip them.arXiv
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildLin et al. · EvaluationEvaluation prompts are mined directly from real anonymized chat logs rather than written by benchmark authors.arXiv
A Careful Examination of Large Language Model Performance on Grade School ArithmeticZhang et al. · EvaluationA fresh unseen grade school math benchmark reveals a measurable overfitting gap versus performance on the original GSM8K.NeurIPS
AlphaMath Almost Zero: Process Supervision without ProcessChen et al. · ReasoningMonte Carlo tree search automatically generates process level supervision for math reasoning without any human annotation.NeurIPS
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsRawles et al. · AgentsA dynamic mobile environment with randomized parameters tests whether agents generalize rather than memorize fixed tasks.arXiv
Aya 23: Open Weight Releases to Further Multilingual ProgressAryabumi et al. · LLMsFocuses instruction tuning compute on twenty three languages to substantially close the multilingual performance gap.arXiv
Chain of Thought Empowers Transformers to Solve Inherently Serial ProblemsLi et al. · ReasoningFormal analysis shows chain of thought lets a fixed depth transformer solve problems that require serial computation.ICLR
Contextual Position Encoding: Learning to Count What's ImportantGolovneva et al. · EfficiencyPositions are computed conditionally on context content rather than token index, letting models better track abstract structure.arXiv
Data-Juicer: A One-Stop Data Processing System for Large Language ModelsChen et al. · SystemsAn operator based data processing system standardizes filtering, deduplication, and mixing pipelines for LLM pretraining.SIGMOD
Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of GeneralizationWang et al. · InterpretabilityExtended training well past memorization, known as grokking, is needed before a transformer generalizes multi-step reasoning.NeurIPS
HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsGutierrez et al. · RAG & MemoryA knowledge graph indexing scheme inspired by the hippocampus lets retrieval integrate information across many documents.NeurIPS
Iterative Reasoning Preference OptimizationPang et al. · ReasoningPreference pairs built from correct versus incorrect chain of thought reasoning traces are used for iterative DPO training.NeurIPS
Learning to Act without ActionsSchmidt and Jiang · RLA policy is pretrained on unlabeled video alone and later grounded to a small amount of true action labeled data.ICLR
LoRA Learns Less and Forgets LessBiderman et al. · EfficiencyCompared with full fine tuning, low rank adaptation learns new tasks less but also forgets prior abilities less.arXiv
Many-Shot In-Context Learning in Multimodal Foundation ModelsJiang et al. · MultimodalScaling the number of in-context demonstrations to hundreds noticeably improves multimodal foundation model task accuracy.arXiv
Not All Language Model Features Are LinearEngels et al. · InterpretabilityFinds multi-dimensional circular feature structures representing concepts like days of the week, challenging pure linearity.arXiv
Octo: An Open-Source Generalist Robot Policy Scaling StudyGhosh et al. · RLTrains a transformer based generalist robot policy on a large open cross embodiment dataset and studies scaling behavior.RSS
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsKim et al. · EvaluationAn open weight evaluator model is trained to closely match proprietary judge models on fine grained scoring rubrics.EMNLP
Prompt Cache: Modular Attention Reuse for Low-Latency InferenceGim et al. · EfficiencyReusing attention states for common prompt prefixes across requests substantially cuts time to first token.MLSys
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM ServingLin et al. · EfficiencyCo-designing a quantization scheme with the serving system pushes practical throughput gains beyond quantization alone.arXiv
Vidur: A Large-Scale Simulation Framework for LLM InferenceAgrawal et al. · SystemsA validated simulator predicts LLM serving performance under different hardware and scheduling configurations without live deployment.MLSys
YOCO: You Only Cache Once, Decoder-Decoder Architectures for Language ModelsSun et al. · EfficiencyA two-part decoder architecture caches key values only once, cutting memory use for very long context generation.arXiv
A Multimodal Automated Interpretability AgentShaham et al. · InterpretabilityAn agent equipped with tools autonomously designs experiments to explain what individual neurons in a network detect.ICML
AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM AgentsGioacchini et al. · AgentsA modular driver-based framework standardizes benchmarking of agent progress across many existing agent task suites.arXiv
Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming ServicesLiu et al. · SystemsProposes a token streaming smoothness metric and scheduler that better matches how users perceive response quality.arXiv
Best Practices and Lessons Learned on Synthetic Data for Language ModelsLiu et al. · LLMsSurveys practical techniques and pitfalls of using model generated synthetic data across pretraining and post-training stages.arXiv
Chinchilla Scaling: A Replication AttemptBesiroglu et al. · LLMsAn independent reanalysis of the original compute-optimal scaling law data finds the reported fit was not fully reproducible.arXiv
Direct Nash Optimization: Teaching Language Models to Self-Improve with General PreferencesRosset et al. · RLOptimizes directly toward a general preference function rather than assuming preferences come from a fixed Bradley-Terry reward.arXiv
Extending Llama-3's Context Ten-Fold OvernightZhang et al. · EfficiencyA short and inexpensive continued pretraining recipe extends an eight thousand token model to eighty thousand tokens.arXiv
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language ModelsZhang et al. · MultimodalAdds a higher resolution visual encoder and multi granularity training to improve region level grounding accuracy.arXiv
Idefics2: Building a Strong Vision-Language ModelLaurencon et al. · MultimodalA fully open reproduction studies which architecture and data decisions most improve open vision language models.Hugging Face
Improving Dictionary Learning with Gated Sparse AutoencodersRajamanoharan et al. · InterpretabilityA gating mechanism in sparse autoencoders separates feature detection from magnitude estimation, sharpening learned features.arXiv
Layer Skip: Enabling Early Exit Inference and Self-Speculative DecodingElhoushi et al. · EfficiencyTraining with layer dropout and an early exit loss lets a single model self-verify its own speculative predictions.ACL
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic EvaluatorsDubois et al. · EvaluationA statistical correction removes the tendency of automatic evaluators to prefer longer responses regardless of quality.arXiv
Let's Think Dot by Dot: Hidden Computation in Transformer Language ModelsPfau et al. · ReasoningFiller tokens with no semantic content still let transformers perform extra hidden computation on some algorithmic tasks.arXiv
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training StrategiesHu et al. · LLMsA sandbox of small models is used to search training hyperparameters that transfer reliably to larger model sizes.arXiv
Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language ModelsRaposo et al. · EfficiencyTokens can skip a layer's computation entirely, letting the network learn its own dynamic compute budget per token.arXiv
Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy DataTajwar et al. · RLShows that generating suboptimal comparison data from the current policy improves preference tuning more than fixed offline data.ICML
RL for Consistency Models: Faster Reward Guided Text-to-Image GenerationOertell et al. · RLReinforcement learning fine tunes fast few-step consistency image generators directly against human preference reward models.arXiv
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language ModelsReka AI Team · MultimodalPresents a family of multimodal models trained from scratch spanning edge to frontier scale compute budgets.arXiv
Scaling Instructable Agents Across Many Simulated WorldsSIMA Team · AgentsA single agent follows free-form language instructions to complete tasks across a wide variety of unrelated 3D video game worlds.Google DeepMind
Scaling Laws for Data Filtering: Data Curation Cannot Be Compute AgnosticGoyal et al. · LLMsShows the best data filtering aggressiveness depends jointly on both dataset size and available training compute.CVPR
The Instruction Hierarchy: Training LLMs to Prioritize Privileged InstructionsWallace et al. · Alignment & SafetyTrains models to treat system, user, and third-party content with different trust levels, reducing prompt injection susceptibility.arXiv
Command R: Retrieval Augmented Generation at Production ScaleCohere Team · LLMsDescribes a model designed from the ground up for retrieval augmented workflows, citations, and tool use in production.Cohere
Cradle: Empowering Foundation Agents Towards General Computer ControlTan et al. · AgentsAn agent perceives raw screen pixels and controls keyboard and mouse directly to operate general purpose software.arXiv
Evaluating Frontier Models for Dangerous CapabilitiesPhuong et al. · Alignment & SafetyIntroduces a suite of tests covering persuasion, cyber offense, self-proliferation, and situational awareness risks.arXiv
Is Cosine-Similarity of Embeddings Really About Similarity?Steck et al. · RAG & MemoryShows regularization choices during training can make cosine similarity an arbitrary and sometimes misleading similarity measure.WWW
Long-Form Factuality in Large Language ModelsWei et al. · EvaluationAn automated pipeline decomposes long generated answers into individual facts and checks each one against search results.arXiv
MM1: Methods, Analysis and Insights from Multimodal LLM Pre-trainingMcKinzie et al. · MultimodalApple's ablation study finds image resolution and connector design matter more than the choice of image encoder.arXiv
RAFT: Adapting Language Model to Domain Specific RAGZhang et al. · RAG & MemoryTraining with a mix of relevant and distractor documents teaches a model to ignore irrelevant retrieved passages.arXiv
RAGAS: Automated Evaluation of Retrieval Augmented GenerationEs et al. · EvaluationA reference-free framework scores retrieval augmented pipelines on faithfulness, relevance, and context quality without human labels.EACL
RewardBench: Evaluating Reward Models for Language Modeling at ScaleLambert et al. · RLA common evaluation suite compares many published reward models on chat, safety, and reasoning preference judgments.arXiv
Safety Cases: How to Justify the Safety of Advanced AI SystemsClymer et al. · Alignment & SafetyProposes structured argument templates that developers could use to justify deploying an advanced AI system safely.arXiv
Sarathi-Serve: Taming Throughput-Latency Tradeoff in LLM Inference with Chunked PrefillsAgrawal et al. · SystemsSplitting large prefill computations into chunks that interleave with decoding steps reduces tail latency under load.OSDI
ShortGPT: Layers in Large Language Models are More Redundant Than You ExpectMen et al. · EfficiencyA simple layer redundancy metric identifies transformer layers that can be removed with minimal performance loss.arXiv
Simple and Scalable Strategies to Continually Pre-train Large Language ModelsIbrahim et al. · LLMsWarm restarts of learning rate and replay of old data let pretraining continue on new data without full retraining.arXiv
Sparse Feature Circuits: Discovering and Editing Interpretable Causal GraphsMarks et al. · InterpretabilitySparse autoencoder features are wired into causal circuits that explain and let researchers edit a model's behavior.arXiv
The Unreasonable Ineffectiveness of the Deeper LayersGromov et al. · EfficiencyA surprisingly large fraction of a large language model's deeper layers can be removed with only minor performance degradation.arXiv
Unsolvable Problem Detection: Evaluating Trustworthiness of Vision Language ModelsYe et al. · EvaluationTests whether vision language models can recognize and refuse questions about images that have no valid answer.arXiv
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?Drouin et al. · AgentsBenchmarks agents on realistic enterprise software tasks like filing forms and searching internal knowledge bases.arXiv
sDPO: Don't Use Your Data All at OnceKim et al. · RLSplitting preference data into chunks and applying DPO stepwise produces a better aligned reference model at each stage.arXiv
A Survey on Data Selection for Language ModelsAlbalak et al. · LLMsSurveys heuristic, model-based, and optimization-based strategies for choosing which pretraining examples to keep or drop.arXiv
A Survey on Knowledge Distillation of Large Language ModelsXu et al. · EfficiencySurveys methods for compressing large teacher language models into smaller students including skill and data distillation.arXiv
AgentOhana: Design Unified Data and Training Pipeline for Effective Agent LearningZhang et al. · AgentsUnifies heterogeneous agent trajectory datasets into a consistent format for training generalist agent models.arXiv
Aligning Modalities in Vision Large Language Models via Preference Fine-TuningZhou et al. · MultimodalApplies direct preference optimization specifically to reduce hallucination in vision language model image descriptions.arXiv
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical DomainsLabrak et al. · LLMsContinued pretraining of an open base model on biomedical text produces a family of domain specialized models.arXiv
Break the Sequential Dependency of LLM Inference Using Lookahead DecodingFu et al. · EfficiencyGenerating and verifying multiple future tokens in parallel using Jacobi iteration reduces the number of decoding steps.ICML
CLLMs: Consistency Large Language ModelsKou et al. · EfficiencyFine tuning a model to map any point on a Jacobi trajectory to the fixed point speeds up parallel decoding.ICML
Chain-of-Thought Reasoning Without Prompting via Decoding Path SearchWang and Zhou · ReasoningChain of thought paths can be found by altering the decoding procedure alone, without any explicit prompt instruction.arXiv
ChemLLM: A Chemical Large Language ModelZhang et al. · LLMsStructured instruction data covering chemical reactions and properties improves a general model's chemistry reasoning.arXiv
Data Engineering for Scaling Language Models to 128K ContextFu et al. · EfficiencyA small amount of carefully constructed long context continued pretraining data is enough to reliably extend context length.ICML
Datasets for Large Language Models: A Comprehensive SurveyLiu et al. · EvaluationSurveys the landscape of pretraining, fine tuning, and evaluation datasets used across recent large language models.arXiv
Debating with More Persuasive LLMs Leads to More Truthful AnswersKhan et al. · Alignment & SafetyHaving stronger models argue opposing sides lets a weaker judge model reach more accurate conclusions than direct questioning.ICML
Dense X Retrieval: What Retrieval Granularity Should We Use?Chen et al. · RAG & MemoryIndexing text as short atomic propositions rather than passages or sentences improves downstream retrieval accuracy.EMNLP
EVA-CLIP-18B: Scaling CLIP to 18 Billion ParametersSun et al. · MultimodalScaling a contrastive image text encoder to eighteen billion parameters yields consistent zero-shot accuracy gains.arXiv
Genstruct: Generating Instruction-Response Pairs from Raw CorporaNousResearch Team · LLMsA generator model converts unlabeled text passages directly into grounded instruction and response training pairs.arXiv
Grasp Multiple Objects with One Hand via Reinforcement LearningWan et al. · RLA single learned dexterous policy grasps multiple different objects using a five fingered robot hand simultaneously.arXiv
Humanoid Locomotion as Next Token PredictionRadosavovic et al. · RLTreating sensorimotor trajectories like language lets a transformer predict humanoid robot actions autoregressively.arXiv
Interpretability Illusions in the Generalization of Simplified ModelsFriedman et al. · InterpretabilitySimplified proxy models used for interpretability can generalize very differently from the original network they represent.ICML
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV CacheLiu et al. · EfficiencyAsymmetric per channel and per token quantization compresses the key value cache to two bits with little accuracy loss.arXiv
Language Models Represent Beliefs of Self and OthersZhu et al. · InterpretabilityLinear probes find a language model encodes a distinguishable internal representation of its own beliefs versus others' beliefs.ICML
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkChen et al. · EvaluationFinds current multimodal models are much weaker judges of image-grounded responses than they are of text-only responses.ICML
Me-LLaMA: Foundation Large Language Models for Medical ApplicationsXie et al. · LLMsContinued pretraining and instruction tuning on biomedical corpora yields models competitive with closed medical systems.arXiv
MemoryBank: Enhancing Large Language Models with Long-Term MemoryZhong et al. · RAG & MemoryA memory mechanism inspired by human forgetting curves lets a chatbot selectively retain important past interactions.AAAI
Nemotron-4 15B Technical ReportNVIDIA · LLMsNVIDIA details pretraining data curation and architecture choices behind a compact bilingual capable foundation model.arXiv
Nomic Embed: Training a Reproducible Long Context Text EmbedderNussbaum et al. · RAG & MemoryReleases a fully open long-context text embedding model along with its training data and code for reproducibility.arXiv
OmniPred: Language Models as Universal RegressorsSong et al. · LLMsA language model trained on text-serialized experiment logs learns to predict numeric outcomes across many different domains.arXiv
On the Societal Impact of Open Foundation ModelsKapoor et al. · Alignment & SafetyArgues the marginal risk of releasing open model weights should be assessed against already available alternatives.arXiv
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksTseng et al. · EfficiencyCombining incoherence processing with lattice vector quantization pushes weight-only compression close to two bits per weight.ICML
Same Task, More Tokens: The Impact of Input Length on the Reasoning Performance of Large Language ModelsLevy et al. · EvaluationHolding task difficulty fixed while only increasing irrelevant input length still measurably degrades reasoning accuracy.ACL
Scaling Laws for Downstream Task Performance in Machine TranslationIsik et al. · LLMsStudies when pretraining data scale reliably predicts translation quality versus when the relationship becomes noisy.ICLR
Simple, Scalable and Effective Clustering for Large-Scale DeduplicationAbbas et al. · LLMsA semantic clustering based deduplication method removes near duplicate pretraining text more effectively than hashing.arXiv
TravelPlanner: A Benchmark for Real-World Planning with Language AgentsXie et al. · AgentsA realistic multi-constraint trip planning benchmark reveals current agents struggle badly with complex constraint satisfaction.ICML
Vision-Language Models as a Source of RewardsBaumli et al. · RLA pretrained vision language model's alignment score to a goal description is used directly as a reward for a control policy.arXiv
Watermarking Makes Language Models RadioactiveSander et al. · Alignment & SafetyText generated by a watermarked model leaves a detectable trace even after it is used to train another model.NeurIPS
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM AgentsMa et al. · EvaluationProvides fine grained progress rate metrics rather than pass or fail outcomes for evaluating multi turn agent tasks.arXiv
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic AgentsAhn et al. · RLA vision language model orchestrates and proposes tasks for a fleet of real robots operating in diverse everyday environments.arXiv
CRUXEval: A Benchmark for Code Reasoning, Understanding, and ExecutionGu et al. · EvaluationTests whether models can predict a short Python function's output or input rather than just generating new code.ICML
Can LLM-Generated Misinformation Be Detected?Chen and Shu · Alignment & SafetyFinds machine generated misinformation can be harder for both humans and detectors to identify than human written misinformation.ICLR
Corrective Retrieval Augmented GenerationYan et al. · RAG & MemoryA lightweight retrieval evaluator triggers web search correction whenever retrieved documents look unreliable or irrelevant.arXiv
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-InferenceHolmes et al. · SystemsCombines dynamic splitfuse scheduling with an optimized inference engine to raise effective generation throughput.arXiv
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM ServingZhong et al. · SystemsRunning prefill and decoding phases on separate GPU pools avoids interference and improves latency service level goals.OSDI
Fairness in Serving Large Language ModelsSheng et al. · SystemsA virtual token counter based scheduler ensures fair throughput allocation among many clients sharing one LLM server.OSDI
Grokking as a First Order Phase Transition in Two Layer NetworksKumar et al. · InterpretabilityA statistical mechanics analysis explains the sudden generalization jump known as grokking as a first order phase transition.ICLR
Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationLiu et al. · EvaluationIntroduces EvalPlus, showing many code benchmark test suites are too weak to catch subtly incorrect generated programs.NeurIPS
LLaVA-NeXT: Improved Reasoning, OCR, and World KnowledgeLiu et al. · MultimodalHigher input resolution and better visual instruction data substantially improve document reading and reasoning ability.arXiv
LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsChen et al. · EfficiencyShifted sparse attention during fine tuning extends context length cheaply while approximating full attention behavior.ICLR
MedPrompt: Can Generalist Foundation Models Outcompete Special-Purpose Tuning?Nori et al. · ReasoningCareful prompting alone lets a general purpose model match specially fine tuned models on medical benchmark questions.arXiv
MimicGen: A Data Generation System for Scalable Robot Learning using Human DemonstrationsMandlekar et al. · RLAutomatically generates large numbers of new robot demonstrations by adapting a small set of human collected trajectories.CoRL
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsGhandeharioun et al. · InterpretabilityReinserting a hidden representation into a different prompt context provides a flexible way to decode what it encodes.ICML
Self-Extend LLM Context Window Without TuningJin et al. · EfficiencyRemapping positional encodings at inference time extends a model's usable context length without any additional training.arXiv
The Impact of Reasoning Step Length on Large Language ModelsJin et al. · ReasoningArtificially lengthening reasoning chains while keeping content fixed still measurably improves downstream task accuracy.ACL
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksKoh et al. · AgentsExtends web agent benchmarks to visually rich tasks that require grounding instructions in rendered page screenshots.arXiv
WARM: On the Benefits of Weight Averaged Reward ModelsRame et al. · RLAveraging the weights of several independently trained reward models reduces reward hacking during RLHF fine tuning.ICML

2023

163 entries
Demonstrate-Search-Predict: Composing Retrieval and Language Models for Knowledge-Intensive NLPKhattab et al. · RAG & MemoryA programming framework composes retrieval and language model calls into pipelines optimized jointly for a task.arXiv
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMsSchulhoff et al. · Alignment & SafetyA large public prompt hacking competition produces a dataset of thousands of real adversarial prompt injection attempts.EMNLP
Textual Adversarial Purification of Large Language ModelsZhang et al. · Alignment & SafetyA purification defense removes adversarial suffixes from a prompt before it reaches the target model for generation.arXiv
Theory of Mind for Multi-Agent Collaboration via Large Language ModelsLi et al. · AgentsModeling teammates' beliefs and intentions explicitly improves cooperative task performance among language model agents.EMNLP
Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyMehrotra et al. · Alignment & SafetyAn automated tree search refines adversarial prompts iteratively until a black box model produces a harmful response.arXiv
Calibrated Language Models Must HallucinateKalai and Vempala · EvaluationA theoretical argument shows any well calibrated language model is mathematically forced to hallucinate some fraction of facts.arXiv
Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language ModelsYu et al. · RAG & MemoryThe model writes sequential reading notes for retrieved documents to judge relevance before producing a final answer.arXiv
Contrastive Chain-of-Thought PromptingChia et al. · ReasoningProviding both valid and invalid reasoning demonstrations as contrast helps models avoid common reasoning mistakes.arXiv
Fine-tuning Language Models for FactualityTian et al. · Alignment & SafetyPreference pairs built automatically from a factuality metric fine-tune a model to be measurably more truthful without human labels.arXiv
GLaMM: Pixel Grounding Large Multimodal ModelRasheed et al. · MultimodalThe first model to produce natural language responses grounded with pixel level segmentation masks for objects mentioned.arXiv
JARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language ModelsWang et al. · AgentsA multimodal memory of past plans lets a Minecraft agent progressively improve at long horizon crafting tasks.arXiv
S-LoRA: Serving Thousands of Concurrent LoRA AdaptersSheng et al. · SystemsUnified paging for adapter weights lets a single server scale to thousands of concurrently served LoRA models.arXiv
Splitwise: Efficient Generative LLM Inference Using Phase SplittingPatel et al. · SystemsSeparating the compute heavy prompt phase from the memory bound generation phase across machines improves cluster throughput.arXiv
The Falcon Series of Open Language ModelsAlmazrouei et al. · LLMsTechnology Innovation Institute describes training and data curation behind the Falcon family of open weight models.arXiv
Universal Jailbreak Backdoors from Poisoned Human FeedbackRando and Tramer · Alignment & SafetyA poisoned reward model can implant a universal backdoor trigger that bypasses RLHF safety training entirely.arXiv
Yuan 2.0: A Large Language Model with Localized Filtering-based AttentionWu et al. · LLMsIEIT Systems proposes a localized attention filter to improve local dependency modeling in a large bilingual model.arXiv
Analogical Prompting: Large Language Models as Analogical ReasonersYasunaga et al. · ReasoningModels self generate relevant exemplars and knowledge before solving a problem instead of relying on fixed demonstrations.arXiv
AutoMix: Automatically Mixing Language ModelsMadaan et al. · EfficiencyA cheap model self-verifies its own answer and escalates only uncertain cases to a more expensive larger model.arXiv
Automatic Model Selection with Large Language Models for Reasoning StrategiesXu et al. · ReasoningA router model picks between several reasoning strategies per question instead of committing to one strategy globally.EMNLP Findings
Branch-Solve-Merge Improves Large Language Model Evaluation and GenerationSaha et al. · ReasoningDecomposing a complex generation or evaluation task into independent branches solved separately and then merged improves quality.arXiv
Contrastive Preference Learning: Learning from Human Feedback without Reinforcement LearningHejna et al. · RLDerives a policy directly from preference data without ever fitting an explicit reward model or running RL.arXiv
Copy Suppression: Comprehensively Understanding an Attention HeadMcDougall et al. · InterpretabilityA single attention head in GPT-2 small is shown to suppress naive token copying to calibrate model confidence.arXiv
Democratizing Reasoning Ability: Tailored Learning from Large Language ModelWang et al. · ReasoningA smaller student model is tailored with interactive multi round learning from a larger teacher's reasoning traces.EMNLP
Detecting Pretraining Data from Large Language ModelsShi et al. · EvaluationA simple statistic based on minimum token probabilities detects whether a given text was likely part of a model's training set.arXiv
FreshQA: Evaluating Large Language Models Against Changing FactsVu et al. · EvaluationA dynamically updated question set exposes how quickly language model knowledge becomes outdated after training.arXiv
Function Vectors in Large Language ModelsTodd et al. · InterpretabilityA compact vector extracted from in context demonstrations can trigger the same task in a fresh context.arXiv
In-Context Learning Creates Task VectorsHendel et al. · InterpretabilityA single vector extracted mid-forward-pass from a demonstration prompt can substitute for the demonstrations themselves.EMNLP Findings
Jailbreaking Black Box Large Language Models in Twenty QueriesChao et al. · Alignment & SafetyAn attacker language model iteratively refines jailbreak prompts against a target using only a handful of queries.arXiv
Large Language Models Cannot Self-Correct Reasoning YetHuang et al. · ReasoningFinds that without external feedback, models often degrade rather than improve their own correct reasoning during self correction.arXiv
MemWalker: Interactive and Long-Context Memory for LLM AgentsChen et al. · RAG & MemoryLong documents are compressed into a navigable tree of summaries that an agent interactively walks to answer queries.arXiv
Multilingual Jailbreak Challenges in Large Language ModelsDeng et al. · Alignment & SafetySafety training that works well in English transfers poorly to lower resource languages, leaving them more exploitable.arXiv
Punica: Multi-Tenant LoRA ServingChen et al. · EfficiencyA custom batched GPU kernel lets many users share one base model while serving distinct LoRA adapters cheaply.arXiv
QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter ModelsFrantar and Alistarh · EfficiencyA custom compression format squeezes trillion parameter mixture of experts models below one bit per parameter.arXiv
RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective AugmentationXu et al. · RAG & MemoryRetrieved documents are compressed into short summaries before augmentation, cutting cost while preserving useful evidence.arXiv
Retrieval meets Long Context Large Language ModelsXu et al. · RAG & MemoryCompares retrieval augmentation against simply extending context length and finds combining both approaches works best.arXiv
Reward Model Ensembles Help Mitigate OveroptimizationCoste et al. · RLAveraging or taking the worst case across an ensemble of reward models reduces reward hacking during policy optimization.arXiv
Self-Consistency for Open-Ended GenerationsJain et al. · ReasoningExtends self consistency sampling and voting beyond fixed answer questions to free form open ended generation tasks.arXiv
Skywork: A More Open Bilingual Foundation ModelWei et al. · LLMsDescribes data contamination detection methods and pretraining recipe for a Chinese and English bilingual base model.arXiv
Specific versus General Principles for Constitutional AIKundu et al. · Alignment & SafetyTests whether a single general good behavior principle can substitute for many narrow constitutional rules.Anthropic
Towards Understanding Sycophancy in Language ModelsSharma et al. · Alignment & SafetyHuman and preference model feedback both reward agreement with a user, revealing a broad sycophancy tendency.Anthropic
Woodpecker: Hallucination Correction for Multimodal Large Language ModelsYin et al. · MultimodalA training free pipeline detects and edits hallucinated objects in vision language model captions using visual grounding.arXiv
Aligning Large Multimodal Models with Factually Augmented RLHFSun et al. · Alignment & SafetyAugments RLHF reward modeling for vision language models with factual image caption data to reduce hallucination.arXiv
Baseline Defenses for Adversarial Attacks Against Aligned Language ModelsJain et al. · Alignment & SafetyEvaluates simple filtering, perturbation, and retokenization defenses against optimization based adversarial prompt attacks.arXiv
CM3Leon: Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction TuningYu et al. · MultimodalA retrieval-augmented autoregressive model trained on licensed image-text data generates and understands images efficiently.arXiv
Certifying LLM Safety against Adversarial PromptingKumar et al. · Alignment & SafetyAn erase-and-check procedure gives a provable guarantee against adversarial suffixes up to a certified attack length.arXiv
Chain-of-Verification Reduces Hallucination in Large Language ModelsDhuliawala et al. · ReasoningModels draft an answer, generate verification questions, answer them independently, and revise the original response accordingly.arXiv
Contrastive Decoding Improves Reasoning in Large Language ModelsO'Brien and Lewis · ReasoningContrasting token probabilities between a strong and weak model at decoding time improves arithmetic and commonsense reasoning.arXiv
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsChuang et al. · Alignment & SafetyContrasting predictions from early and late transformer layers during decoding reduces factual hallucination without retraining.arXiv
GPTFuzzer: Red Teaming Large Language Models with Auto-Generated Jailbreak PromptsYu et al. · Alignment & SafetyA mutation based fuzzing framework automatically generates and evolves jailbreak prompts against target chat models.arXiv
Petals: Collaborative Inference and Fine-tuning of Large ModelsBorzunov et al. · SystemsVolunteers pool consumer GPUs over the internet to jointly run inference and fine tuning of huge open models.arXiv
Query Rewriting for Retrieval-Augmented Large Language ModelsMa et al. · RAG & MemoryA small trainable rewriter reformulates the user query before retrieval to better match the underlying reader model.arXiv
Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4Guo et al. · AgentsAn agent reasons about what opponents believe to make stronger decisions in imperfect information card games.arXiv
Text2Reward: Reward Shaping with Language Models for Reinforcement LearningXie et al. · RLA language model writes and iteratively refines dense reward code for reinforcement learning tasks from natural language.arXiv
Textbooks Are All You Need II: phi-1.5 technical reportLi et al. · LLMsMicrosoft researchers show a 1.3B model trained on synthetic textbook style data reasons surprisingly well for its size.arXiv
The Reversal Curse: LLMs Trained on A is B Fail to Learn B is ABerglund et al. · InterpretabilityModels trained only on statements in one direction fail to generalize and answer correctly when the relation is reversed.arXiv
AgentSims: An Open-Source Sandbox for Large Language Model EvaluationLin et al. · AgentsA configurable simulated town lets researchers evaluate specific agent capabilities in controllable everyday scenarios.arXiv
Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language ModelsSel et al. · ReasoningA single prompt teaches the model to imitate algorithmic search behavior rather than expanding many separate branches.arXiv
Cumulative Reasoning with Large Language ModelsZhang et al. · ReasoningIntermediate conclusions accumulate as reusable propositions, letting later reasoning steps build directly on earlier verified results.arXiv
Language Reward Modulation for Pretraining Reinforcement LearningAdeniji et al. · RLA vision language model's alignment score is used as a pretraining reward signal before any task specific reward exists.arXiv
Linearity of Relation Decoding in Transformer Language ModelsHernandez et al. · InterpretabilityMany factual relations are approximated by a single linear transformation applied to the subject token representation.arXiv
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language ModelsAwadalla et al. · MultimodalAn open reproduction of Flamingo lets the research community study and extend few-shot vision language models freely.arXiv
SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step ReasoningMiao et al. · ReasoningA model checks each step of its own solution without an external answer key, catching many of its own reasoning errors.arXiv
Simple Synthetic Data Reduces Sycophancy in Large Language ModelsWei et al. · Alignment & SafetyAdding synthetic examples where correct answers disagree with a stated user opinion measurably reduces sycophantic behavior.arXiv
Studying Large Language Model Generalization with Influence FunctionsGrosse et al. · InterpretabilityScales influence functions to billion parameter models to trace which training examples shaped a given generation.Anthropic
ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive StudyYao et al. · EfficiencySystematically studies how weight and activation quantization choices interact across many large language model sizes.arXiv
Building Cooperative Embodied Agents Modularly with Large Language ModelsZhang et al. · AgentsSeparate perception, memory, communication, and planning modules built from language models let embodied agents cooperate.arXiv
Distilling Reasoning Capabilities into Smaller Language ModelsShridhar et al. · ReasoningDecomposes a teacher's chain of thought into subproblems so a small student model can learn to reason step by step.ACL Findings
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsYe et al. · EvaluationDecomposes evaluation into twelve fine grained skills to give a more interpretable picture than one overall score.arXiv
Focused Transformer: Contrastive Training for Context ScalingTworkowski et al. · EfficiencyA contrastive training objective teaches a memory augmented transformer to better distinguish relevant from irrelevant context.NeurIPS
How Is ChatGPT's Behavior Changing Over Time?Chen et al. · EvaluationRepeated identical queries to the same chatbot API over several months reveal substantial unannounced drift in its answers.arXiv
In-context Autoencoder for Context Compression in a Large Language ModelGe et al. · EfficiencyCompresses long context into a small number of memory slots that a frozen language model can still condition on.arXiv
Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsTrivedi et al. · RAG & MemoryRetrieval and chain of thought generation are interleaved step by step so each new fact guides the next query.ACL
Large Language Models as General Pattern MachinesMirchandani et al. · ReasoningPretrained language models extrapolate abstract sequence patterns well enough to guide simple robot control tasks.CoRL
MMBench: Is Your Multi-modal Model an All-around Player?Liu et al. · EvaluationA hierarchical multimodal benchmark with circular evaluation strategy improves robustness of multiple choice grading.arXiv
Measuring Faithfulness in Chain-of-Thought ReasoningLanham et al. · InterpretabilityPerturbing chain of thought text and observing answer changes reveals how much models actually rely on it.Anthropic
Parallel Context Windows for Large Language ModelsRatner et al. · EfficiencySplits a long input into separately encoded windows whose attentions are merged, extending usable context without retraining.ACL
Question Decomposition Improves the Faithfulness of Model-Generated ReasoningRadhakrishnan et al. · InterpretabilityBreaking a question into subquestions before answering produces reasoning traces that better reflect the true computation.Anthropic
SEED-Bench: Benchmarking Multimodal Large Language ModelsLi et al. · EvaluationA large scale multiple choice benchmark measures both spatial and temporal image and video understanding abilities.arXiv
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel GenerationNing et al. · ReasoningModels first outline an answer skeleton then expand each point in parallel to cut end to end latency.arXiv
XGen-7B Technical Report: Long Context Language ModelingNijkamp et al. · LLMsSalesforce presents a 7B parameter model trained with an eight thousand token context length from the start.arXiv
Are Aligned Neural Networks Adversarially Aligned?Carlini et al. · Alignment & SafetyDemonstrates that current safety alignment techniques provide only weak robustness against optimization based adversarial attacks.NeurIPS
Baichuan-7B: An Open Large-Scale Pre-Trained Language ModelBaichuan Inc. · LLMsDetails the pretraining corpus, tokenizer, and evaluation of an open 7B Chinese and English bilingual base model.arXiv
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT ModelsWang et al. · EvaluationEvaluates GPT models across toxicity, bias, privacy, robustness, and fairness under both benign and adversarial prompts.NeurIPS
Emergent and Predictable Memorization in Large Language ModelsBiderman et al. · InterpretabilityMemorization behavior of a fully trained large model can be predicted in advance from the behavior of smaller partially trained ones.NeurIPS
Fast Segment AnythingZhao et al. · MultimodalReformulating segmentation as detection followed by prompt guided selection runs fifty times faster than the original model.arXiv
Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingWu et al. · RLCollecting feedback at the level of individual text spans rather than whole responses yields more targeted reward signals.NeurIPS
Ghost in the Minecraft: Generally Capable Agents for Open-World EnvironmentsZhu et al. · AgentsDecomposing goals into text based subgoals with a knowledge base lets an agent explore Minecraft far more broadly.arXiv
InternLM: A Multilingual Language Model with Progressively Enhanced CapabilitiesTeam InternLM · LLMsShanghai AI Laboratory introduces a bilingual foundation model trained in stages to improve reasoning and coding skill.arXiv
Language to Rewards for Robotic Skill SynthesisYu et al. · RLA language model translates natural language instructions into reward parameters that a separate optimizer turns into robot motion.CoRL
LeanDojo: Theorem Proving with Retrieval-Augmented Language ModelsYang et al. · ReasoningAn open toolkit and benchmark for retrieval augmented neural theorem proving in the Lean proof assistant.NeurIPS
OpenLLaMA: An Open Reproduction of LLaMAGeng and Liu · LLMsAn open source permissively licensed reproduction of the LLaMA architecture trained on public datasets from scratch.arXiv
Orca-style Explanation Tuning: Progressive Learning from GPT-4 TracesMukherjee et al. · LLMsTraining on rich explanation traces rather than short answers teaches a smaller model deeper reasoning behavior.arXiv
Otter: A Multi-Modal Model with In-Context Instruction TuningLi et al. · MultimodalInstruction tuning on a curated multimodal in context dataset improves an image and video language model's instruction following.arXiv
PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial PromptsZhu et al. · EvaluationA unified benchmark measures how much small adversarial prompt perturbations degrade large language model accuracy.arXiv
Retrieval-Augmented Multimodal Language ModelingYasunaga et al. · MultimodalRetrieving relevant image-text pairs from an external memory during generation reduces the parameters needed to store world knowledge.ICML
Segment Anything in High QualityKe et al. · MultimodalA lightweight token added to the frozen Segment Anything model produces noticeably higher quality mask boundaries.NeurIPS
Shikra: Unleashing Multimodal LLM's Referential Dialogue MagicChen et al. · MultimodalThe model reads and outputs spatial coordinates directly in natural language to support referring and grounding dialogue.arXiv
ToolQA: A Dataset for LLM Question Answering with External ToolsZhuang et al. · EvaluationA benchmark isolates whether correct answers came from genuine tool use rather than memorized parametric knowledge.NeurIPS
Towards Measuring the Representation of Subjective Global Opinions in Language ModelsDurmus et al. · Alignment & SafetyIntroduces a cross country survey benchmark showing model responses skew toward opinions from wealthier Western countries.Anthropic
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video UnderstandingZhang et al. · MultimodalSeparate video and audio branches are aligned to a language model so it can discuss both modalities together.arXiv
Visual Adversarial Examples Jailbreak Aligned Large Language ModelsQi et al. · Alignment & SafetyA single adversarially crafted image can be optimized to jailbreak an otherwise safety-aligned vision language model.arXiv
Adapting Language Models to Compress ContextsChevalier et al. · EfficiencyTrains a model to recursively compress earlier text into summary vectors, extending effective context far beyond the window.arXiv
Adversarial Demonstration Attacks on Large Language ModelsWang et al. · Alignment & SafetyPoisoned in context demonstrations can silently steer a model's predictions without altering its underlying weights.arXiv
AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackDubois et al. · RLA low cost simulated annotator lets researchers rapidly iterate on RLHF methods without expensive human labeling loops.NeurIPS
Automatic Prompt Optimization with Gradient Descent and Beam SearchPryzant et al. · ReasoningTreats prompt edits like textual gradients and uses beam search to iteratively improve a prompt's performance.arXiv
Blockwise Parallel Transformers for Large Context ModelsLiu and Abbeel · EfficiencyFuses feedforward and attention computation blockwise to reduce memory overhead for very long context transformers.arXiv
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingGou et al. · ReasoningModels critique their own outputs using external tool feedback and then revise their answer in a closed loop.arXiv
Do Large Language Models Know What They Don't Know?Yin et al. · EvaluationIntroduces a benchmark of unanswerable questions to test whether models can recognize the limits of their own knowledge.ACL Findings
Efficiently Scaling Transformer InferencePope et al. · SystemsAnalyzes partitioning strategies across chips to minimize latency and cost when serving huge transformer models.MLSys
Evaluating Object Hallucination in Large Vision-Language ModelsLi et al. · EvaluationIntroduces a polling based evaluation method that reliably measures how often vision language models describe absent objects.EMNLP
Faith and Fate: Limits of Transformers on CompositionalityDziri et al. · ReasoningTransformers solve compositional tasks like multiplication by pattern matching subgraphs rather than by true systematic reasoning.NeurIPS
FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving PerformanceChen et al. · EfficiencyCascading cheap models before expensive ones and caching prior answers cuts API cost substantially with little accuracy loss.arXiv
Jailbreaking ChatGPT via Prompt Engineering: An Empirical StudyLiu et al. · Alignment & SafetyCategorizes and empirically tests a taxonomy of manually crafted jailbreak prompt patterns collected from online communities.arXiv
Landmark Attention: Random-Access Infinite Context Length for TransformersMohtashami and Jaggi · EfficiencySpecial landmark tokens let a transformer randomly access far away context blocks without attending to all of them.arXiv
Meta-in-context learning in large language modelsCoda-Forno et al. · ReasoningRepeated in context learning episodes recursively improve a model's own few shot learning ability within a session.NeurIPS
PaLI-X: On Scaling Up a Multilingual Vision and Language ModelChen et al. · MultimodalScaling both the vision and language components together improves a wide range of multilingual vision language benchmarks.arXiv
PandaGPT: One Model to Instruction-Follow Them AllSu et al. · MultimodalCombining a multimodal encoder with an instruction tuned language model gives one system six modalities of input.arXiv
Prompting Is Not a Substitute for Probability Measurements in Large Language ModelsHu and Levy · EvaluationShows that prompting a model for a probability estimate diverges from its actual internal probability computed directly.EMNLP
Self-Polish: Enhance Reasoning in Large Language Models via Problem RefinementXi et al. · ReasoningThe model rewrites a confusing problem statement into clearer form before attempting to solve it step by step.arXiv
Symbol Tuning Improves In-Context Learning in Language ModelsWei et al. · ReasoningReplacing natural language labels with arbitrary symbols during tuning forces models to rely more on context.arXiv
Tab-CoT: Zero-shot Tabular Chain of ThoughtJin and Lu · ReasoningReformats chain of thought reasoning as a table so intermediate steps stay structured and easier to verify.ACL Findings
The Curse of Recursion: Training on Generated Data Makes Models ForgetShumailov et al. · Alignment & SafetyRepeatedly training models on their own generated data causes progressive collapse of the tails of the data distribution.arXiv
The False Promise of Imitating Proprietary LLMsGudibande et al. · EvaluationShows imitation models fine tuned on outputs from a stronger model look good on style but not on factuality.arXiv
VideoChat: Chat-Centric Video UnderstandingLi et al. · MultimodalCombines a video foundation model with a language model and a dialogue interface for interactive video question answering.arXiv
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign LanguagesChen et al. · MultimodalTreats image, speech, and video as distinct foreign languages that separate interfaces translate into a language model's space.arXiv
Boosting Theory-of-Mind Performance in Large Language Models via PromptingMoghaddam and Honey · ReasoningSimple few shot and chain of thought prompts substantially raise GPT-4 performance on false belief reasoning tasks.arXiv
Evaluating Verifiability in Generative Search EnginesLiu et al. · EvaluationHuman evaluation of AI powered search engines finds a large share of generated statements lack a supporting citation.EMNLP Findings
Instruction Tuning with GPT-4Peng et al. · LLMsReleases an instruction dataset generated by GPT-4 and shows it improves instruction following over GPT-3.5 data.arXiv
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction ModelGao et al. · MultimodalAdds a small number of trainable parameters to LLaMA so it gains visual instruction following without full fine tuning.arXiv
Learning to Compress Prompts with Gist TokensMu et al. · EfficiencyA model learns to compress a long instruction prompt into a handful of gist tokens that a frozen model can condition on instead.NeurIPS
Localizing Model Behavior with Path PatchingGoldowsky-Dill et al. · InterpretabilityIntroduces path patching, a causal intervention technique for isolating which attention paths drive a network's behavior.arXiv
The Internal State of an LLM Knows When It's LyingAzaria and Mitchell · InterpretabilityA simple classifier trained on hidden layer activations can detect whether a generated statement is true fairly reliably.EMNLP Findings
Towards Automated Circuit Discovery for Mechanistic InterpretabilityConmy et al. · InterpretabilityAn automated algorithm reconstructs the sparse circuits inside a transformer that previously required manual discovery.NeurIPS
mPLUG-Owl: Modularization Empowers Large Language Models with MultimodalityYe et al. · MultimodalA modular training approach aligns a visual encoder to a frozen language model without harming its text abilities.arXiv
Anthropic's Core Views on AI Safety: When, Why, What, and HowAnthropic · Alignment & SafetyLays out the reasoning behind Anthropic's founding strategy of engaging with frontier AI development to improve its safety.Anthropic
Can AI-Generated Text Be Reliably Detected?Sadasivan et al. · EvaluationArgues that paraphrasing attacks can reliably evade essentially all current statistical and watermark-based text detectors.arXiv
Eliciting Latent Predictions from Transformers with the Tuned LensBelrose et al. · InterpretabilityA learned affine probe per layer decodes intermediate hidden states into vocabulary predictions more reliably than logit lens.arXiv
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUSheng et al. · SystemsAn offloading and scheduling strategy across GPU, CPU, and disk memory maximizes batch throughput on one accelerator.ICML
GPT-4 Passes the Bar ExamKatz et al. · EvaluationReports GPT-4 scoring near the top decile of real bar exam takers, a large jump over the previous model generation.Philosophical Transactions of the Royal Society A
Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied AgentsHuang et al. · AgentsCombines a language model's plans with grounded affordance functions so a robot only proposes physically feasible actions.NeurIPS
Language Models can Solve Computer TasksKim et al. · AgentsA recursive critique and improve prompting loop lets a language model complete multi step computer interface tasks.arXiv
Large Language Models Are Human-Level Prompt EngineersZhou et al. · ReasoningAn automatic prompt engineering method searches over candidate instructions and selects the one scoring best on a task.ICLR
Larger Language Models Do In-Context Learning DifferentlyWei et al. · ReasoningLarger models can override semantic priors and learn from flipped or unusual labels shown only in context.arXiv
Query2doc: Query Expansion with Large Language ModelsWang et al. · RAG & MemoryA model generates a hypothetical passage to expand a short search query before it is passed to a retrieval system.EMNLP
Whose Opinions Do Language Models Reflect?Santurkar et al. · Alignment & SafetyUses public opinion surveys to show model outputs align more closely with some demographic groups than others.arXiv
Active Prompting with Chain-of-Thought for Large Language ModelsDiao et al. · ReasoningUses model uncertainty to select which questions most need human annotated chain of thought exemplars.arXiv
Aligning Language Models with Preferences through f-divergence MinimizationGo et al. · RLGeneralizes RLHF objectives to a family of f-divergences, unifying several existing preference alignment methods as special cases.arXiv
AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingLi et al. · SystemsStatistically multiplexing several large models across shared GPUs improves overall serving latency under bursty load.OSDI
Contrastive Search Is What You Need for Neural Text GenerationSu and Collier · ReasoningA degeneration penalty during decoding based on token similarity to prior context produces more coherent open-ended text.TMLR
Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal ProofsJiang et al. · ReasoningA model drafts an informal proof, sketches it into a formal outline, and fills gaps with an automated prover.ICLR
Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security AttacksKang et al. · Alignment & SafetyShows classic security techniques like payload splitting and obfuscation reliably bypass content filters on instruction-tuned models.arXiv
Guiding Pretraining in Reinforcement Learning with Large Language ModelsDu et al. · RLA language model suggests useful exploration goals that shape an intrinsic reward signal for a reinforcement learning agent.ICML
Multimodal Chain-of-Thought Reasoning in Language ModelsZhang et al. · MultimodalA two-stage framework first generates a rationale grounded in an image and then uses it to produce the final answer.arXiv
Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt InjectionGreshake et al. · Alignment & SafetyShows attackers can hide instructions in a webpage or document that hijack an LLM application retrieving that content.arXiv
Poisoning Web-Scale Training Datasets Is PracticalCarlini et al. · Alignment & SafetyDemonstrates that an attacker with a modest budget could poison a meaningful fraction of common web-scraped training datasets.arXiv
Pretraining Language Models with Human PreferencesKorbak et al. · Alignment & SafetyCompares objectives for injecting human preference signals directly during pretraining rather than only at fine tuning.arXiv
The Capacity for Moral Self-Correction in Large Language ModelsGanguli et al. · Alignment & SafetyLarger instruction tuned models can follow a simple instruction to avoid stereotype and bias driven outputs.Anthropic
Batch Prompting: Efficient Inference with Large Language Model APIsCheng et al. · EfficiencyGrouping several inputs into a single prompt cuts token and cost overhead while mostly preserving per example accuracy.arXiv
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskLi et al. · InterpretabilityA GPT trained only on Othello move sequences forms an internal board state representation of the game.ICLR
Generate Rather than Retrieve: Large Language Models Are Strong Context GeneratorsYu et al. · RAG & MemoryReplacing document retrieval with a language model that generates its own contextual passage directly improves open-domain QA.ICLR
Recitation-Augmented Language ModelsSun et al. · RAG & MemoryModels first recite relevant knowledge from their own parameters before answering, mimicking retrieval without external documents.ICLR
Secrets of RLHF in Large Language Models Part II: Reward ModelingWang et al. · RLA companion study analyzes reward model quality, measurement, and ensembling strategies that stabilize RLHF training.arXiv
Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningCreswell et al. · ReasoningSplitting reasoning into an alternating selection and inference module produces more faithful and accurate multi step answers.ICLR