<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>HuggingFace Papers — Trending</title>
<link>https://cybertar.youngood.tech/huggingface-papers.html</link>
<description>Trending papers on HuggingFace, curated by Cybertar.</description>
<language>en</language>
<generator>Cybertar RSS Generator</generator>
<lastBuildDate>Mon, 22 Jun 2026 11:12:06 +0000</lastBuildDate>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<atom:link href="https://cybertar.youngood.tech/feeds/rss/huggingface-papers-trending.xml" rel="self" type="application/rss+xml" />
<item>
<title>A decoder-only foundation model for time-series forecasting</title>
<link>https://huggingface.co/papers/2310.10688</link>
<guid isPermaLink="true">https://huggingface.co/papers/2310.10688</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>A large language model adapted for time-series forecasting achieves near-optimal zero-shot performance on diverse datasets across different time scales and granularities. 4 authors · Published on Oct 14, 2023 Upvote 34 GitHub 24.9k arXiv Pa</p>]]></description>
</item>
<item>
<title>GLM-5: from Vibe Coding to Agentic Engineering</title>
<link>https://huggingface.co/papers/2602.15763</link>
<guid isPermaLink="true">https://huggingface.co/papers/2602.15763</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by taesiri  GLM-5 advances foundation models with DSA for cost reduction, asynchronous reinforcement learning for improved alignment, and enhanced coding capabilities for real-world software engineering. 186 authors · Published on</p>]]></description>
</item>
<item>
<title>TradingAgents: Multi-Agents LLM Financial Trading Framework</title>
<link>https://huggingface.co/papers/2412.20138</link>
<guid isPermaLink="true">https://huggingface.co/papers/2412.20138</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>A multi-agent framework using large language models for stock trading simulates real-world trading firms, improving performance metrics like cumulative returns and Sharpe ratio. 4 authors · Published on Dec 28, 2024 Upvote 101 GitHub 87.9k </p>]]></description>
</item>
<item>
<title>SkillOpt: Executive Strategy for Self-Evolving Agent Skills</title>
<link>https://huggingface.co/papers/2605.23904</link>
<guid isPermaLink="true">https://huggingface.co/papers/2605.23904</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by taesiri  SkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multip</p>]]></description>
</item>
<item>
<title>EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning</title>
<link>https://huggingface.co/papers/2601.02163</link>
<guid isPermaLink="true">https://huggingface.co/papers/2601.02163</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>EverMemOS presents a self-organizing memory system for large language models that processes dialogue streams into structured memory cells and scenes to enhance long-term interaction capabilities. 11 authors · Published on Jan 5, 2026 Upvote</p>]]></description>
</item>
<item>
<title>PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training</title>
<link>https://huggingface.co/papers/2606.03264</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.03264</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by ChengCui  PaddleOCR-VL-1.6 enhances document parsing performance through targeted data optimization and progressive post-training techniques, achieving state-of-the-art results on OmniDocBench v1.6. PaddlePaddle · Published on </p>]]></description>
</item>
<item>
<title>OpenDevin: An Open Platform for AI Software Developers as Generalist
  Agents</title>
<link>https://huggingface.co/papers/2407.16741</link>
<guid isPermaLink="true">https://huggingface.co/papers/2407.16741</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by akhaliq  OpenDevin is a platform for developing AI agents that interact with the world by writing code, using command lines, and browsing the web, with support for multiple agents and evaluation benchmarks. 24 authors · Publish</p>]]></description>
</item>
<item>
<title>Kronos: A Foundation Model for the Language of Financial Markets</title>
<link>https://huggingface.co/papers/2508.02739</link>
<guid isPermaLink="true">https://huggingface.co/papers/2508.02739</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Kronos, a specialized pre-training framework for financial K-line data, outperforms existing models in forecasting and synthetic data generation through a unique tokenizer and autoregressive pre-training on a large dataset. 7 authors · Publ</p>]]></description>
</item>
<item>
<title>FastContext: Training Efficient Repository Explorer for Coding Agents</title>
<link>https://huggingface.co/papers/2606.14066</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.14066</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by qiushao  FastContext separates repository exploration from code solving in LLM agents using specialized exploration models that reduce token consumption and improve resolution rates. Microsoft · Published on Jun 12, 2026 Upvote</p>]]></description>
</item>
<item>
<title>MinerU2.5: A Decoupled Vision-Language Model for Efficient
  High-Resolution Document Parsing</title>
<link>https://huggingface.co/papers/2509.22186</link>
<guid isPermaLink="true">https://huggingface.co/papers/2509.22186</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by taesiri  MinerU2.5, a 1.2B-parameter document parsing vision-language model, achieves state-of-the-art recognition accuracy with computational efficiency through a coarse-to-fine parsing strategy. 61 authors · Published on Sep </p>]]></description>
</item>
<item>
<title>Efficient Memory Management for Large Language Model Serving with
  PagedAttention</title>
<link>https://huggingface.co/papers/2309.06180</link>
<guid isPermaLink="true">https://huggingface.co/papers/2309.06180</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by akhaliq  PagedAttention algorithm and vLLM system enhance the throughput of large language models by efficiently managing memory and reducing waste in the key-value cache. 9 authors · Published on Sep 12, 2023 Upvote 59 GitHub </p>]]></description>
</item>
<item>
<title>VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models</title>
<link>https://huggingface.co/papers/2606.16140</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.16140</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by SenXu1123  VibeThinker-3B demonstrates that compact models can achieve state-of-the-art performance on verifiable reasoning tasks through specialized training techniques, challenging conventional scaling assumptions. WeiboAI · </p>]]></description>
</item>
<item>
<title>LTX-2: Efficient Joint Audio-Visual Foundation Model</title>
<link>https://huggingface.co/papers/2601.03233</link>
<guid isPermaLink="true">https://huggingface.co/papers/2601.03233</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by taesiri  LTX-2 is an open-source audiovisual diffusion model that generates synchronized video and audio content using a dual-stream transformer architecture with cross-modal attention and classifier-free guidance. 29 authors ·</p>]]></description>
</item>
<item>
<title>LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference</title>
<link>https://huggingface.co/papers/2510.09665</link>
<guid isPermaLink="true">https://huggingface.co/papers/2510.09665</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>LMCACHE enables efficient KV cache management for large language models by storing caches outside GPU memory, supporting cache reuse across queries and inference engines while achieving significant throughput improvements. 11 authors · Publ</p>]]></description>
</item>
<item>
<title>Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory</title>
<link>https://huggingface.co/papers/2504.19413</link>
<guid isPermaLink="true">https://huggingface.co/papers/2504.19413</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by akhaliq  Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory syste</p>]]></description>
</item>
<item>
<title>Toward Generalist Autonomous Research via Hypothesis-Tree Refinement</title>
<link>https://huggingface.co/papers/2606.11926</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.11926</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by namespace-ERI  An AI framework called Arbor enables autonomous scientific research by combining strategic coordination, isolated hypothesis testing, and a persistent knowledge tree to iteratively improve research outcomes acros</p>]]></description>
</item>
<item>
<title>JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence</title>
<link>https://huggingface.co/papers/2606.14777</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.14777</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by iieycx  A vision-language model operates continuously in real-time, making autonomous decisions about when to respond or delegate, enabling interactive systems that perceive and act upon environmental changes without user promp</p>]]></description>
</item>
<item>
<title>ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration</title>
<link>https://huggingface.co/papers/2605.03042</link>
<guid isPermaLink="true">https://huggingface.co/papers/2605.03042</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by RuofengYang  ARIS is an open-source research harness that uses cross-model adversarial collaboration to ensure reliable long-term research outcomes through coordinated execution, orchestration, and assurance layers. Shanghai Ji</p>]]></description>
</item>
<item>
<title>Recursive Language Models</title>
<link>https://huggingface.co/papers/2512.24601</link>
<guid isPermaLink="true">https://huggingface.co/papers/2512.24601</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by rajkumarrawal  We study allowing large language models (LLMs) to process arbitrarily long prompts through the lens of inference-time scaling. We propose  (RLMs), a general inference strategy that treats long prompts as part of </p>]]></description>
</item>
<item>
<title>SmolDocling: An ultra-compact vision-language model for end-to-end
  multi-modal document conversion</title>
<link>https://huggingface.co/papers/2503.11576</link>
<guid isPermaLink="true">https://huggingface.co/papers/2503.11576</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by andito  SmolDocling is a compact vision-language model that performs end-to-end document conversion with robust performance across various document types using 256M parameters and a new markup format. IBM Granite · Published on</p>]]></description>
</item>
<item>
<title>OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation</title>
<link>https://huggingface.co/papers/2410.17799</link>
<guid isPermaLink="true">https://huggingface.co/papers/2410.17799</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>A novel GPT-based model, OmniFlatten, enables real-time natural full-duplex spoken dialogue through a multi-stage post-training technique that integrates speech and text without altering the original model's architecture. 9 authors · Publis</p>]]></description>
</item>
<item>
<title>DreamX-World 1.0: A General-Purpose Interactive World Model</title>
<link>https://huggingface.co/papers/2606.16993</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.16993</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by AdinaY  DreamX-World 1.0 is a interactive text/image-to-video model that generates long-horizon content with camera control and scene persistence using specialized encoding, training techniques, and optimization methods. AMAP-M</p>]]></description>
</item>
<item>
<title>MOSS-TTS Technical Report</title>
<link>https://huggingface.co/papers/2603.18090</link>
<guid isPermaLink="true">https://huggingface.co/papers/2603.18090</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by fdugyt  MOSS-TTS is a speech generation model using discrete audio tokens and autoregressive modeling with capabilities for voice cloning, pronunciation control, and long-form generation across multiple languages. OpenMOSS · Pu</p>]]></description>
</item>
<item>
<title>Recursive Multi-Agent Systems</title>
<link>https://huggingface.co/papers/2604.25917</link>
<guid isPermaLink="true">https://huggingface.co/papers/2604.25917</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by jiaruz2  RecursiveMAS extends recursive scaling principles from single models to multi-agent systems, enabling collaborative reasoning through iterative latent-space computations with improved efficiency and accuracy. Stanford </p>]]></description>
</item>
<item>
<title>Cosmos 3: Omnimodal World Models for Physical AI</title>
<link>https://huggingface.co/papers/2606.02800</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.02800</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by taesiri  Cosmos 3 is an omnimodal world model that processes and generates multiple data types through a unified mixture-of-transformers architecture, achieving state-of-the-art performance in various understanding and generati</p>]]></description>
</item>
<item>
<title>RF-DETR: Neural Architecture Search for Real-Time Detection Transformers</title>
<link>https://huggingface.co/papers/2511.09554</link>
<guid isPermaLink="true">https://huggingface.co/papers/2511.09554</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by mervenoyan  RF-DETR, a light-weight detection transformer, uses weight-sharing NAS to optimize accuracy and latency for real-time detection across diverse datasets. Roboflow · Published on Nov 12, 2025 Upvote 12 GitHub 8.08k ar</p>]]></description>
</item>
<item>
<title>Multi-module GRPO: Composing Policy Gradients and Prompt Optimization
  for Language Model Programs</title>
<link>https://huggingface.co/papers/2508.04660</link>
<guid isPermaLink="true">https://huggingface.co/papers/2508.04660</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>mmGRPO, a multi-module extension of GRPO, enhances accuracy in modular AI systems by optimizing LM calls and prompts across various tasks. 13 authors · Published on Aug 6, 2025 Upvote 3 GitHub 35.3k arXiv Page  mmGRPO, a multi-module extens</p>]]></description>
</item>
<item>
<title>DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI</title>
<link>https://huggingface.co/papers/2512.16676</link>
<guid isPermaLink="true">https://huggingface.co/papers/2512.16676</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by zbhpku  DataFlow is an LLM-driven data preparation framework that enhances data quality and reproducibility for various tasks, improving LLM performance with automatically generated pipelines. Peking University · Published on D</p>]]></description>
</item>
<item>
<title>Zep: A Temporal Knowledge Graph Architecture for Agent Memory</title>
<link>https://huggingface.co/papers/2501.13956</link>
<guid isPermaLink="true">https://huggingface.co/papers/2501.13956</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Zep, a memory layer service, outperforms MemGPT in the DMR benchmark and LongMemEval by excelling in dynamic knowledge integration and temporal reasoning, critical for enterprise use cases. 5 authors · Published on Jan 20, 2025 Upvote 12 Gi</p>]]></description>
</item>
<item>
<title>Kairos: A Native World Model Stack for Physical AI</title>
<link>https://huggingface.co/papers/2606.16533</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.16533</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by shanyou92  Kairos is a world model framework that learns from diverse experiences, maintains persistent states through hybrid temporal attention mechanisms, and operates efficiently across different hardware platforms for physi</p>]]></description>
</item>
<item>
<title>SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning</title>
<link>https://huggingface.co/papers/2606.10804</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.10804</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by taesiri  SCAIL-2 enables end-to-end character animation by directly transferring motion from driving videos without intermediate representations, using unified task decomposition and synthetic data generation. Z.ai · Published </p>]]></description>
</item>
<item>
<title>Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models</title>
<link>https://huggingface.co/papers/2606.03748</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.03748</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by nielsr  YOLO26 addresses real-time vision challenges through a unified model family with NMS-free inference, improved training strategies, and multi-task capabilities spanning detection, segmentation, and pose estimation. Ultra</p>]]></description>
</item>
<item>
<title>COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation</title>
<link>https://huggingface.co/papers/2605.31264</link>
<guid isPermaLink="true">https://huggingface.co/papers/2605.31264</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by jasonrqh  Person-grounded AI skills are automatically distilled from heterogeneous traces into inspectable, correctable packages that capture both capabilities and behavioral patterns. shanghai ailab · Published on May 29, 2026</p>]]></description>
</item>
<item>
<title>Very Large-Scale Multi-Agent Simulation in AgentScope</title>
<link>https://huggingface.co/papers/2407.17789</link>
<guid isPermaLink="true">https://huggingface.co/papers/2407.17789</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by akhaliq  Enhancements to the AgentScope platform improve scalability, efficiency, and ease of use for large-scale multi-agent simulations through distributed mechanisms, flexible environments, and user-friendly tools. 8 authors</p>]]></description>
</item>
<item>
<title>AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets</title>
<link>https://huggingface.co/papers/2512.10971</link>
<guid isPermaLink="true">https://huggingface.co/papers/2512.10971</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>AI-Trader presents the first fully automated live benchmark for evaluating large language models in financial decision-making across multiple markets with autonomous information processing. 6 authors · Published on Dec 1, 2025 Upvote 9 GitH</p>]]></description>
</item>
<item>
<title>Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents</title>
<link>https://huggingface.co/papers/2604.22085</link>
<guid isPermaLink="true">https://huggingface.co/papers/2604.22085</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by MoeinAbtahi  Memanto presents a universal memory layer for agentic AI that eliminates computational overhead of hybrid semantic graph architectures through a typed semantic memory schema and information-theoretic search engine.</p>]]></description>
</item>
<item>
<title>LightRAG: Simple and Fast Retrieval-Augmented Generation</title>
<link>https://huggingface.co/papers/2410.05779</link>
<guid isPermaLink="true">https://huggingface.co/papers/2410.05779</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>LightRAG improves Retrieval-Augmented Generation by integrating graph structures for enhanced contextual awareness and efficient information retrieval, achieving better accuracy and response times. 5 authors · Published on Oct 8, 2024 Upvot</p>]]></description>
</item>
<item>
<title>AgentScope 1.0: A Developer-Centric Framework for Building Agentic
  Applications</title>
<link>https://huggingface.co/papers/2508.16279</link>
<guid isPermaLink="true">https://huggingface.co/papers/2508.16279</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by taesiri  AgentScope enhances agentic applications by providing flexible tool-based interactions, unified interfaces, and advanced infrastructure based on the ReAct paradigm, supporting efficient and safe development and deploym</p>]]></description>
</item>
<item>
<title>PDFMathTranslate: Scientific Document Translation Preserving Layouts</title>
<link>https://huggingface.co/papers/2507.03009</link>
<guid isPermaLink="true">https://huggingface.co/papers/2507.03009</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>PDFMathTranslate enables layout-preserving scientific document translation using large language models and precise layout detection, offering improved precision, flexibility, and efficiency. 4 authors · Published on Jul 2, 2025 Upvote 2 Git</p>]]></description>
</item>
<item>
<title>VibeVoice Technical Report</title>
<link>https://huggingface.co/papers/2508.19205</link>
<guid isPermaLink="true">https://huggingface.co/papers/2508.19205</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by unilm  VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity. Microsoft Research · Published on Aug 26, 2</p>]]></description>
</item>
<item>
<title>SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture</title>
<link>https://huggingface.co/papers/2605.12500</link>
<guid isPermaLink="true">https://huggingface.co/papers/2605.12500</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by Paranioar  Unified vision-language models treat understanding and generation as integrated processes rather than separate tasks, demonstrating strong performance across multiple multimodal capabilities including image synthesis</p>]]></description>
</item>
<item>
<title>IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot
  Text-To-Speech System</title>
<link>https://huggingface.co/papers/2502.05512</link>
<guid isPermaLink="true">https://huggingface.co/papers/2502.05512</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>IndexTTS, an enhanced text-to-speech system combining XTTS and Tortoise models, offers improved naturalness, enhanced voice cloning, and controllable usage through hybrid character-pinyin modeling and optimized vector quantization. 5 author</p>]]></description>
</item>
<item>
<title>PyTorch Distributed: Experiences on Accelerating Data Parallel Training</title>
<link>https://huggingface.co/papers/2006.15704</link>
<guid isPermaLink="true">https://huggingface.co/papers/2006.15704</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>The PyTorch distributed data parallel module optimizes large-scale model training using techniques like gradient bucketing, computation-communication overlap, and selective synchronization to achieve near-linear scalability. 11 authors · Pu</p>]]></description>
</item>
<item>
<title>Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of
  Encoders</title>
<link>https://huggingface.co/papers/2408.15998</link>
<guid isPermaLink="true">https://huggingface.co/papers/2408.15998</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by akhaliq  Mixture of vision encoders and resolutions in multimodal large language models improves performance through concatenation of visual tokens and a Pre-Alignment mechanism, leading to superior results on benchmarks. 15 au</p>]]></description>
</item>
<item>
<title>Agent READMEs: An Empirical Study of Context Files for Agentic Coding</title>
<link>https://huggingface.co/papers/2511.12884</link>
<guid isPermaLink="true">https://huggingface.co/papers/2511.12884</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by hao-li  Agentic coding tools receive goals written in natural language as input, break them down into specific tasks, and write or execute the actual code with minimal human intervention. Central to this process are agent conte</p>]]></description>
</item>
<item>
<title>SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning</title>
<link>https://huggingface.co/papers/2606.13673</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.13673</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by cmhungsteve  SpatialClaw is a training-free framework that uses code as an action interface to enable flexible, stateful spatial reasoning in vision-language models, achieving superior performance across diverse 3D/4D spatial r</p>]]></description>
</item>
<item>
<title>Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?</title>
<link>https://huggingface.co/papers/2606.08063</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.08063</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by Jiaqi-hkust  Robust-U1 enhances multimodal large language models' robustness against visual corruptions through self-recovery capabilities that improve both visual quality and reasoning performance. 9 authors · Published on Jun</p>]]></description>
</item>
<item>
<title>SIA: Self Improving AI with Harness &amp; Weight Updates</title>
<link>https://huggingface.co/papers/2605.27276</link>
<guid isPermaLink="true">https://huggingface.co/papers/2605.27276</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by VigneshHexo  A self-improving AI framework simultaneously updates both model weights and task-specific agent architecture through a language-model feedback agent across legal classification, GPU optimization, and biological dat</p>]]></description>
</item>
<item>
<title>Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance</title>
<link>https://huggingface.co/papers/2606.19195</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.19195</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by Uyoung  A lightweight image inpainting framework achieves high-fidelity results with significantly reduced parameters and inference time through novel local-global interaction blocks and adaptive distillation strategies. 6 auth</p>]]></description>
</item>
<item>
<title>Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games</title>
<link>https://huggingface.co/papers/2606.19338</link>
<guid isPermaLink="true">https://huggingface.co/papers/2606.19338</guid>
<pubDate>Mon, 22 Jun 2026 11:11:58 +0000</pubDate>
<description><![CDATA[<p><strong>Authors:</strong> Unknown</p><p>Submitted by ChrisDing1105  A new benchmark suite called RNG-Bench is introduced to evaluate multimodal foundation models' ability to reconstruct past observations and use them for decision-making in multi-step interactions, featuring two g</p>]]></description>
</item>
</channel>
</rss>
