HuggingFaceAlaya-EVOKE: Endless World Generation Research Paper OverviewAug 14, 20261 minThe Alaya-EVOKE paper discusses moving from linear-scaling supervision to endless world generation. It has gained 64 upvotes on HuggingFace.
ArXivOptimizing Meta-Harnesses for Long-Horizon Agentic DesignAug 14, 20261 minA new research paper introduces AutoDesign, a method for optimizing meta-harnesses to improve long-horizon agentic design. The paper received 28 upvotes on HuggingFace.
ArXivAI4AI Test-Time Strong-to-Weak Capability Transfer via HarnessesAug 13, 20261 minA new research paper titled 'AI4AI at Test-Time' has been published, focusing on strong-to-weak capability transfer via harnesses. It received 71 upvotes on Hugging Face.
HuggingFaceSci-VBench: Benchmarking Scientific Video Generation ReasoningAug 11, 20261 minA new benchmark called Sci-VBench has been created to evaluate video generation models on scientifically rigorous and reasoning-heavy tasks. It has received 21 upvotes.
HuggingFaceExtracting Reasoning Traces from Proprietary LLM APIsAug 11, 20261 minA new paper demonstrates a method to extract internal reasoning traces from closed-source, proprietary LLM APIs. The research has received 4 upvotes on HuggingFace.
HuggingFaceMacaron-V1: Self-Improving Continual Learning with Mixture-of-LoRAAug 11, 20261 minA new research paper introduces Macaron-V1, an approach for open continual learning that uses self-improvement techniques with Mixture-of-LoRA. The paper has gained 46 upvotes on HuggingFace.
HuggingFaceSWE-Bench ProMax: Benchmark for Large-Scale Multilingual Code RefactoringAug 11, 20261 minA new benchmark, SWE-Bench ProMax, has been introduced to evaluate agents on large-scale, multilingual code refactoring tasks. It has received 74 upvotes on HuggingFace.
ArXivThe Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic DistillationJul 28, 20261 minA research paper proposing a framework for training AI agents to perform multi-turn, long-horizon planning through on-policy distillation from single and multiple teacher models. The work explores pre-training and post-training paradigms.
ArXivThe Regression Tax: Decomposing why skills help and hurt LLM agentsJul 27, 20261 minAn ArXiv paper analyzes how learned skills can both benefit and hinder LLM agents, introducing the concept of a 'regression tax' that explains performance trade-offs.
HuggingFaceScaling Native Multimodal Pre-Training From ScratchJul 27, 20261 minA research paper explores training a multimodal model from scratch using a unified architecture, achieving strong performance across vision-language tasks. It received 12 upvotes on Hugging Face.