<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>arXiv AIGC 图像/视频生成论文精选</title>
    <description>自动筛选 arXiv 上 AIGC 图像生成、视频生成、扩散模型相关的最新论文</description>
    <link>https://arxiv.org</link>
    <atom:link href="http://101.32.50.67:8080/arxiv_aigc_feed.xml" rel="self" type="application/rss+xml"/>
    <lastBuildDate>Sat, 19 Sep 2026 09:00:28 GMT</lastBuildDate>
    <language>zh-cn</language>
    <generator>arXiv AIGC RSS Generator v1.0</generator>
    <item>
      <title>[多模态生成] [模型架构] Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision</title>
      <link>https://arxiv.org/abs/2609.20820v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20820v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:59:53 GMT</pubDate>
      <dc:creator>Nitish Dashora, Douglas Chen, Idan Shenfeld et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades perf...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Nitish Dashora, Douglas Chen, Idan Shenfeld et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many approaches to policy memory involve compressing historical information through expensive VLM queries in-the-loop to process only task-salient information. In this paper, we propose an alternative approach in which computationally intensive VLM queries are made during train-time to learn a lightweight latent memory that can be efficiently queried at deployment time. Our representation, which we call the \textbf{workspace token}, is trained by (1) using a VLM to identify current and historical information necessary for completing a task, then (2) distilling these into the workspace token using a set-reconstruction decoder loss. In both simulation and hardware, we show that the workspace token can be used as a drop-in replacement for observations during deployment, enabling policies to solve memory-intensive tasks without the need for VLM reasoning in-the-loop. Interestingly, we found that workspace tokens are not only more lightweight but also lead to better policy performance.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20820v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20820v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[控制与编辑] SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos</title>
      <link>https://arxiv.org/abs/2609.20818v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20818v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:59:41 GMT</pubDate>
      <dc:creator>Peiyu Liu, Dingxi Zhang, Federico Tombari et al.</dc:creator>
      <category>控制与编辑</category>
      <description>A splash lives for a fraction of a second: sheets tear into ligaments and droplets, appearance is view-dependent and nearly textureless, and little persists long enough to track. Reconstruction resear...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Peiyu Liu, Dingxi Zhang, Federico Tombari et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 控制与编辑</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> style transfer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">A splash lives for a fraction of a second: sheets tear into ligaments and droplets, appearance is view-dependent and nearly textureless, and little persists long enough to track. Reconstruction research has consequently focused on smoke, synthetic liquids, or gently deforming surfaces. To our knowledge, no synchronized multi-view dataset of splashing liquids exists. We therefore introduce a benchmark of 20 real scenes, from coherent streams to violent splashes, captured by seven synchronized, calibrated 4K cameras at 60 fps, with manually refined per-view liquid and container masks and fixed evaluation splits. We further present SplashSplat, built on a single principle: impose physical structure only where the observations can constrain it. Per-frame liquid SDFs fused from the masks provide the geometry, level-set transport between consecutive SDFs yields a coarse velocity field, and Lagrangian carriers advected along this flow, corrected against each new observation and reseeded where coverage is lost, decode local Gaussians for differentiable rendering. SplashSplat outperforms state-of-the-art dynamic Gaussian splatting methods on our real captures and on a synthetic benchmark, with physically more plausible motion and a lower training cost. The same representation supports temporal interpolation and style transfer without re-optimization.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20818v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20818v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 控制与编辑 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations</title>
      <link>https://arxiv.org/abs/2609.20817v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20817v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:59:40 GMT</pubDate>
      <dc:creator>Kevin Qu, Tao Sun, Massimiliano Viola et al.</dc:creator>
      <category>模型架构</category>
      <description>Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kevin Qu, Tao Sun, Massimiliano Viola et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our model jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. We further propose an observed articulation span objective that supervises the motion range each part exhibits across the input observations, encouraging the model to leverage the full observation set. To overcome the limited scale and diversity of existing datasets, we introduce a procedural data generator that synthesizes self-annotated assets during training. Experiments on PartNet-Mobility, ACD, and ArtiCraft-10K demonstrate consistent improvements over both feed-forward and optimization-based baselines. Project page: https://kevinqu7.github.io/famos</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20817v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20817v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Paint-Anything: Unified Any-Color Control for Image Generation and Editing</title>
      <link>https://arxiv.org/abs/2609.20816v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20816v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:59:30 GMT</pubDate>
      <dc:creator>Ji Xie, Dewei Zhou, Xinyu Huang et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Professional design requires any-color control: the ability to specify an object&apos;s target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, ed...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Paint-Anything: Unified Any-Color Control for Image Generation and Editing</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ji Xie, Dewei Zhou, Xinyu Huang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> t2i, image generation, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Professional design requires any-color control: the ability to specify an object&apos;s target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, editing, and colorization, but often relies on dedicated color representations or specialized inference procedures. Advances in large language models offer a simpler starting point: even compact models can associate hex values with color semantics. We present Paint-Anything, which learns a shared hex-prompt interface for generation and editing through object-level color supervision. We develop a data pipeline that constructs Paint-500K from real images through object grounding, perceptual color labeling, and editing-pair synthesis. Since shadows make real-image labels only approximate colors, we complement this supervision with pure-color anchors whose pixels exactly match their paired hex values. These anchors are used only at high-noise timesteps, leaving low-noise training to natural images. We further introduce Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks. On FLUX.2-4B, Paint-Anything improves ACBench-T2I and ACBench-Edit scores by 85.3% and 28.3%, respectively, relative to the base model, with ablations supporting the training recipe. It also achieves the highest average CompColor score among the compared methods.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20816v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20816v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis</title>
      <link>https://arxiv.org/abs/2609.20815v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20815v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:59:28 GMT</pubDate>
      <dc:creator>Zahra Ghaffari, Massih Bahar, Mojgan Forootan et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a broad spectrum of extracolonic tumors. Early identification and accurate classification of these ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zahra Ghaffari, Massih Bahar, Mojgan Forootan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a broad spectrum of extracolonic tumors. Early identification and accurate classification of these syndromes are essential for timely diagnosis, individualized patient management, and targeted surveillance strategies for affected families. However, public endoscopic datasets are largely organized around the individual sporadic polyp, and none links the polyposis phenotype to histopathology and germline findings at the patient level. Here, we present ERCPMP-Gx, an endoscopic, histopathological, and genomic dataset developed to support the application of artificial intelligence (AI) in the recognition, characterization, and classification of colorectal polyposis. Most procedures were performed using the Olympus EVIS X1 system with white-light endoscopy (WLE), narrow-band imaging (NBI), magnifying NBI (M-NBI), and NBI with near focus modes, yielding 160 images and accompanying video clips. Approximately eighty percent of cases represent clinically and/or genetically confirmed hereditary polyposis syndromes (PG), including familial adenomatous polyposis (FAP), Peutz-Jeghers syndrome (PJS), juvenile polyposis syndrome (JPS), and ganglioneuroma syndrome (GNS), while the remaining twenty percent comprise non-hereditary polyps and polyp-mimicking lesions with overlapping morphological features (Non-PG), included to support differential classification. Each released record is linked, where available, to standardized endoscopic annotations, representative histopathology, and clinically reported germline findings, forming an AI-ready, patient-level annotation framework. The dataset is publicly accessible at Mendeley (https://doi.org/10.17632/nzyfc544bx.2). For the latest updates and further information, readers are referred to the DataBioX website: https://databiox.com.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20815v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20815v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Score Centering Stabilizes Off-policy Reinforcement Learning</title>
      <link>https://arxiv.org/abs/2609.20807v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20807v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:58:17 GMT</pubDate>
      <dc:creator>Martin Marek, Max Ryabinin</dc:creator>
      <category>模型架构</category>
      <description>Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). H...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Score Centering Stabilizes Off-policy Reinforcement Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Martin Marek, Max Ryabinin</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely eliminating TIM is impractical, as it would come at a major cost to rollout efficiency. In this paper, we show that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step. We derive an additive &quot;score centering&quot; correction term that stabilizes RL under TIM by canceling drift. When training models from 0.6B to 30B parameters, score centering alone matches or outperforms methods based on importance sampling under quantization, with the gap growing as the mismatch becomes more severe. Because the correction is additive, score centering also composes with importance sampling -- their composition outperforms pure importance-sampling baselines in our staleness experiments.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20807v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20807v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning</title>
      <link>https://arxiv.org/abs/2609.20784v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20784v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:52:08 GMT</pubDate>
      <dc:creator>Yan Yu, Zhengxi Lu, Yizhou Liu et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yan Yu, Zhengxi Lu, Yizhou Liu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher&apos;s success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20784v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20784v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations</title>
      <link>https://arxiv.org/abs/2609.20779v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20779v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:49:28 GMT</pubDate>
      <dc:creator>Sarah Wyer, Sue Black, Noura Al Moubayed</dc:creator>
      <category>模型架构</category>
      <description>Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically in...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sarah Wyer, Sue Black, Noura Al Moubayed</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men&apos;s rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36\% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($ρ= +0.55$, $p = .034$) while Detoxify does not ($ρ= -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20779v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20779v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies</title>
      <link>https://arxiv.org/abs/2609.20776v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20776v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:48:07 GMT</pubDate>
      <dc:creator>Xin Chen, Sen Chen, Yujuan Ding et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different ta...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xin Chen, Sen Chen, Yujuan Ding et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> flow matching, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf{GeoAAC}, a geometry-based adaptive action chunking method for flow-based VLA policies that adjusts the action horizon according to the reliability of the current action prediction. We show that the geometry of Flow Matching denoising trajectories provides process-level information for characterizing prediction reliability, with geometric variation across action prefixes remaining positively correlated with predictive uncertainty. GeoAAC uses this prefix-wise geometry to construct a horizon-wise geometric profile and adaptively determine the action horizon from a single generation without additional training. Experiments with GR00T N1.5 and π0.5 on LIBERO, LIBERO-Pro, RoboCasa365, and real-world manipulation tasks show consistent improvements over fixed-action-horizon baselines and existing adaptive methods, including up to 8.7 percentage points in simulation and an increase in average real-world success rate from 53.3\% to 74.4\%.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20776v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20776v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants</title>
      <link>https://arxiv.org/abs/2609.20769v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20769v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:46:07 GMT</pubDate>
      <dc:creator>Tianao Li, Xinhui Qian, Emma Alexander</dc:creator>
      <category>扩散模型</category>
      <description>Flow matching has emerged as the state-of-the-art generative model and has been used for plug-and-play (PnP) priors to solve inverse problems in computational imaging. However, existing flow-based inv...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tianao Li, Xinhui Qian, Emma Alexander</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> flow matching, diffusion</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Flow matching has emerged as the state-of-the-art generative model and has been used for plug-and-play (PnP) priors to solve inverse problems in computational imaging. However, existing flow-based inverse solvers assume linear forward models and/or make simplifying approximations in posterior sampling. To circumvent these problems, we introduce FlowSGS, a flow-based posterior sampling method using Split Gibbs Sampling (SGS) to decompose the posterior into a likelihood step and a prior step. Specifically, we sample from the likelihood step using Langevin dynamics and leverage the Stochastic Interpolants (SI) framework to integrate a pretrained flow model into the prior step. We provide a form for the prior step that uses SI&apos;s reverse-time SDE, and show connections to previous PnP methods. Moreover, with the aid of the flow prior&apos;s straight probability paths and a novel timestep correction technique for the reverse-time SDE, FlowSGS requires fewer network evaluations in its prior step than plug-and-play diffusion samplers. Our experiments show state-of-the-art performance on a range of inverse problems. For the first time, we provide an experiment on a nonlinear inverse problem (Fourier phase retrieval) for flow-based inverse solvers.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20769v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20769v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols</title>
      <link>https://arxiv.org/abs/2609.20765v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20765v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:44:22 GMT</pubDate>
      <dc:creator>Tariq Abdul-Quddoos, Xiangfang Li, Lijun Qian</dc:creator>
      <category>模型架构</category>
      <description>Radio Frequency(RF)-Fingerprinting is a spectrum monitoring technique that identifies specific transmitters based on hardware impairments imprinted within the emitted signal. Although widely researche...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tariq Abdul-Quddoos, Xiangfang Li, Lijun Qian</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Radio Frequency(RF)-Fingerprinting is a spectrum monitoring technique that identifies specific transmitters based on hardware impairments imprinted within the emitted signal. Although widely researched, studies almost exclusively consider scenarios where only one transmitter is emitting at a time, limiting real world applicability. In this work, we further the study of RF-Fingerprinting by considering co-channel interference, with multiple emitted signals interfering with each other, overlapping in time and frequency. Specifically, we formulate this problem as a multi-label classification problem and employ a 1D convolutional neural network (CNN). Furthermore, the models are calibrated such that the confidence thresholds for the label probabilities are derived, with guarantees on the upper bound on the average number of False Negatives, providing a degree of confidence in not missing a true spectrum policy violation. The proposed method is validated using real world data from the POWDER 5G testbed on devices transmitting 802.11a(Wi-Fi), 4G LTE, and 5G NR waveforms. The results show accuracy as high as 97% and as low as 73% after calibration depending on channel conditions. Also calibrating for various average false negatives upper bounds achieves micro recall scores of approximately (1 - calibrated false negatives) with the calibration robust to out-of-distribution interference, demonstrating the potential of the proposed method in a realistic high contention wireless environment</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20765v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20765v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Large Language Models as Falsifiers for Cyber-Physical Systems</title>
      <link>https://arxiv.org/abs/2609.20752v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20752v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:40:04 GMT</pubDate>
      <dc:creator>Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak</dc:creator>
      <category>模型架构</category>
      <description>Falsification searches for counterexamples to formal specifications in cyber-physical systems (CPS). With specifications written in Signal Temporal Logic (STL), falsification can be formulated as a ro...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Large Language Models as Falsifiers for Cyber-Physical Systems</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Falsification searches for counterexamples to formal specifications in cyber-physical systems (CPS). With specifications written in Signal Temporal Logic (STL), falsification can be formulated as a robustness optimization problem, traditionally tackled with black-box search algorithms. In parallel, large language models (LLMs) have recently emerged as surprisingly effective optimizers when coupled with iterative prompting. In this work, we connect these ideas and introduce LLM-Falsifier, an LLM-based approach that falsifies specifications by minimizing the STL robustness degree. Beyond generic prompt-based optimization, our key idea is to expose the LLM to semantic information that is natural for language models but absent from standard numerical optimizers, including natural-language input and output names, output trajectories, and critical-time witnesses for the minimum robustness value. These additions enable smarter and more sample-efficient robustness search. On the ARCH-COMP falsification benchmarks, LLM-Falsifier is shown to outperform existing falsification tools based on a range of optimization paradigms, from surrogate-based and Bayesian optimization to search-based testing, on 14 of 21 specifications when measured by the average number of simulations required to find a counterexample.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20752v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20752v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] dQwen3.5: Hybrid-Attention Diffusion Language Models</title>
      <link>https://arxiv.org/abs/2609.20751v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20751v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:39:46 GMT</pubDate>
      <dc:creator>Anton Xue, Litu Rout, Aditya Akella et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling ha...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">dQwen3.5: Hybrid-Attention Diffusion Language Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Anton Xue, Litu Rout, Aditya Akella et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapting Qwen3.5 at 0.8B, 2B, 4B, and 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation: against a full-attention control, the hybrid reaches a given training loss in about half the tokens. Across scales, dQwen3.5 resembles full-attention DLMs in any-order decoding behavior and performs strongly under parallel decoding.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20751v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20751v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [模型架构] [扩散模型] [图像生成] Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation</title>
      <link>https://arxiv.org/abs/2609.20744v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20744v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:33:15 GMT</pubDate>
      <dc:creator>Haocheng Xi, Yiming Xie, Hexu Zhao et al.</dc:creator>
      <category>视频生成</category>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Haocheng Xi, Yiming Xie, Hexu Zhao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> video generation, distillation, dit, diffusion, diffusion model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which updates memory once per frame by jointly incorporating its spatial tokens. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5x speedup over the 50-step dense H3 baseline on the same GPU count.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20744v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20744v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models</title>
      <link>https://arxiv.org/abs/2609.20722v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20722v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:16:31 GMT</pubDate>
      <dc:creator>Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner et al.</dc:creator>
      <category>模型架构</category>
      <description>Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and ca...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and causal head-level attribution to autonomously discover optimal steering parameters. Across three scales (1B x 3, 2-3B x 2, and 7-9B x 4), our engine achieves 16.7 percentage-point improvement on spam at 1B (standard deviation 4.7; 39 runs), with gains increasing to 21 to 42 percentage points at 7-9B across four architectures. On SST-2 sentiment, it achieves a 13.1 percentage-point improvement with zero code changes. Mechanistic grounding enables automated discovery of intervention points that generalize across tasks and architectures. On sentiment, RepE without head masking fails to improve over baseline, while Deep Noir improves all models (p less than 0.01). We further show that steering creates a predictable prompt-injection attack surface whose vulnerability increases monotonically with steering magnitude. This finding is relevant to agent systems deploying steered classifiers.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20722v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20722v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Don&apos;t Mask the Environment: Observation Supervision Changes How Agents Explore Under RL</title>
      <link>https://arxiv.org/abs/2609.20715v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20715v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:11:29 GMT</pubDate>
      <dc:creator>Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as conte...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Don&apos;t Mask the Environment: Observation Supervision Changes How Agents Explore Under RL</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20715v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20715v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation</title>
      <link>https://arxiv.org/abs/2609.20700v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20700v1</guid>
      <pubDate>Thu, 17 Sep 2026 17:02:53 GMT</pubDate>
      <dc:creator>Lili Wang, Jing Li, Xiaowen Sun et al.</dc:creator>
      <category>模型架构</category>
      <description>Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, wit...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Lili Wang, Jing Li, Xiaowen Sun et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> u-net, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-vendor cardiac MRI the mean $Δ$Dice from adaptation is statistically indistinguishable from zero while 58.7% of cases are individually made worse. We quantify this harm as harmful accepted area (HA), the harmful fraction of the edited area a controller deploys. Held-out tuning gives a stronger baseline than a fixed horizon, but the budget it selects transfers on neither of the two main medical benchmarks, and no global budget can condition on the case. We show that prediction fragmentation---the disagreement geometry between $M_0$ and the adapted mask $M_k$---predicts HA with no labels or extra backward passes at decision time, comparably on three benchmarks (Spearman $ρ$ 0.50--0.60), at a quarter of gradient-norm&apos;s latency. A case-level router built on it cuts HA from 0.228 to 0.139 on a benchmark that took no part in its design, with the design frozen and only cut-points recalibrated there. On the cardiac benchmark the design was selected on, the router cuts HA from 0.129 to 0.013 at matched Dice and 1.10 deployed updates, against the retrospective-best budget found post hoc on evaluation labels, and reduces that 58.7% to 20.0%, an upper bound we quantify. Where the retained cases are not net-helped (as on prostate), the router still cuts HA but concedes accuracy, a boundary we report. Thresholds are fit once on a labeled split disjoint from evaluation; decisions use no labels or gradients. The template ports across architecture and domain (nnU-Net$\to$SegFormer, Cityscapes$\to$ACDC) with coordinate, thresholds and per-bucket actions instantiated per domain.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20700v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20700v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] TetrisCNN for interpretable detection of phases of matter from experimental quantum simulator data</title>
      <link>https://arxiv.org/abs/2609.20693v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20693v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:58:52 GMT</pubDate>
      <dc:creator>Kacper Cybiński, Björn van Zwol, James Enouen et al.</dc:creator>
      <category>模型架构</category>
      <description>Detecting phases of matter in general relies on identifying the correct order parameter - a task that remains notoriously difficult for unknown transitions and traditionally is guided by physical intu...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">TetrisCNN for interpretable detection of phases of matter from experimental quantum simulator data</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kacper Cybiński, Björn van Zwol, James Enouen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Detecting phases of matter in general relies on identifying the correct order parameter - a task that remains notoriously difficult for unknown transitions and traditionally is guided by physical intuition and educated guess. Neural networks have recently offered an alternative route by locating phase transitions in known models without any a priori physical knowledge. Yet these approaches remain black boxes and only identify phases without elucidating their properties. Moreover, they often struggle when confronted with realistic, noisy experimental data, which constitute the ultimate testbed for automated methods in physics. Here, we bridge these perspectives by introducing TetrisCNN, a convolutional architecture with parallel branches of differently shaped filters, reminiscent of Tetris blocks, that learns sparse, interpretable latent representations directly in terms of spin correlators. Applied to experimental snapshots of two-dimensional Ising and XY quantum simulators measured in multiple bases, the network not only detects phase transitions and crossovers but also expresses its latent representation and decision boundaries as symbolic formulas built from experimentally measurable spin correlators. This framework opens the way to integrating interpretable neural networks with quantum simulators to uncover and understand new phases of matter.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20693v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20693v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Epidemiological Causal Graph Identification: Challenges, Identifiability and Algorithms</title>
      <link>https://arxiv.org/abs/2609.20676v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20676v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:48:45 GMT</pubDate>
      <dc:creator>Sambit Mishra, Yingying Wang, Christine K. Johnson et al.</dc:creator>
      <category>模型架构</category>
      <description>Causal discovery from observational data is fundamental to statistics and machine learning, yet determining causal direction without interventions necessitates structural assumptions. Existing identif...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Epidemiological Causal Graph Identification: Challenges, Identifiability and Algorithms</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sambit Mishra, Yingying Wang, Christine K. Johnson et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Causal discovery from observational data is fundamental to statistics and machine learning, yet determining causal direction without interventions necessitates structural assumptions. Existing identifiability research primarily focuses on continuous variables under additive noise models, often neglecting mixed datasets containing ordinal scales, counts, and continuous measurements. This paper investigates causal discovery in Directed Acyclic Graphs (DAGs) where nodes follow either an ordinal distribution (via an ordered logit model) or a regular one-parameter exponential family distribution. We prove that the edge direction between an ordinal and an exponential family node is distributionally identifiable for generic parameter values. Our findings generalize previous Ordinal-Poisson results to the broader exponential family. Computationally, we introduce a score-based exhaustive search and a masked continuous optimization framework using DAGMA for larger graphs. Numerical results validate the theory, recovering edge orientations within a Markov equivalence class that are unidentifiable under classical structural equation models.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20676v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20676v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents</title>
      <link>https://arxiv.org/abs/2609.20673v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20673v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:48:05 GMT</pubDate>
      <dc:creator>Dennis Rotondi, Abdelrhman Werby, Kai O. Arras</dc:creator>
      <category>模型架构</category>
      <description>To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene rep...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Dennis Rotondi, Abdelrhman Werby, Kai O. Arras</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vae</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP_{50} points for movable parts, 2.8 AP_{50} points under joint origin-and-axis constraints, and 6.7 AP_{50} points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20673v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20673v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] [扩散模型] Learning Foresight without Explicit Trajectories for 3D Diffusion Policies</title>
      <link>https://arxiv.org/abs/2609.20669v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20669v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:44:13 GMT</pubDate>
      <dc:creator>Zhongbo Zhang, Zaibin Zhang, Yifan Wang et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also ant...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Learning Foresight without Explicit Trajectories for 3D Diffusion Policies</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhongbo Zhang, Zaibin Zhang, Yifan Wang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, 3d diffusion, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from action learning. We introduce Movement Trend Guidance, a simple but effective way to provide this foresight without introducing an explicit plan. From a short observation history, the policy learns a compact latent representation of interaction evolution. During training, sparse future gripper states supervise this representation; at inference, only the latent is retained as future-oriented conditioning alongside the current observation. The latent provides global conditioning for action generation, while an additional gated FiLM branch is used only at the UNet bottleneck. Despite adding only 3.52% more parameters to DP3, our method preserves the original dense-action and receding-horizon formulation and consistently improves upon DP3 across RoboTwin2.0, LIBERO-40, and DexArt. It reaches 62.8% vs. 56.1% in 50-task RoboTwin2.0 mixed training, 71.93% vs. 37.08% on LIBERO-40, and 72.0% vs. 49.0% on five real-robot tasks. These results show that a diffusion policy can benefit substantially from knowing where an interaction is heading, without being told exactly where to move.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20669v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20669v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Earth Surface Immune System for Rapid Monitoring of Unknown Anomalies</title>
      <link>https://arxiv.org/abs/2609.20662v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20662v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:39:26 GMT</pubDate>
      <dc:creator>Jingtao Li, Qian Zhu, Xinyu Wang et al.</dc:creator>
      <category>模型架构</category>
      <description>Earth surface anomalies, driven by escalating climate change, and expanding human activities, are increasing in both frequency and diversity, yet their limited historical data and unpredictability mak...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Earth Surface Immune System for Rapid Monitoring of Unknown Anomalies</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jingtao Li, Qian Zhu, Xinyu Wang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Earth surface anomalies, driven by escalating climate change, and expanding human activities, are increasing in both frequency and diversity, yet their limited historical data and unpredictability make them fundamentally different from conventional remote sensing targets. Existing methods address specific anomaly categories or stop at localization, leaving a gap between detection and actionable information. Here we present ESIA, an Earth Surface Immune System whose architecture is constrained by three principles from the biological immune system, refined over millions of years against equally diverse and uncertain threats. A non-specific innate immune stage treats anomalies as unobserved changes in time-series satellite imagery, generating binary localization maps at 14.51 km2/s without assuming any anomaly category, surpassing the strongest general baseline by 37% in F1. A specific adaptive immune stage applies negative selection to filter text prompts and matches surviving prompts with localized image patches through a multi-modal foundation model, enabling open-vocabulary recognition of unknown anomaly attributes including category, affected area, and damage severity, with recognition F1 exceeding 80%. A mutation mechanism tunes minimal embeddings at test time, adapting to each scene in 3.26s using a single reference image pair. We validate ESIA on a global-scale dataset covering 19,801.60 km2 across six anomaly categories, comparing against 22 models, and further apply it to quantify degraded farmland in the Dnipro Delta following the Kakhovka Dam collapse and assess burn severity from 2025 Palisades Fire in Los Angeles. This unprecedented flexibility in handling unknown anomalies opens new avenues for real-time disaster response and environmental surveillance.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20662v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20662v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface</title>
      <link>https://arxiv.org/abs/2609.20659v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20659v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:38:37 GMT</pubDate>
      <dc:creator>Zimu Han, Yiming Zeng, Jiyao Zhang et al.</dc:creator>
      <category>模型架构</category>
      <description>Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-spe...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zimu Han, Yiming Zeng, Jiyao Zhang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-training. During handheld UMI demonstrations, HIL-UMI queries the current policy on the same observation stream without executing its predictions. The Energy Score compares the human action trajectory with policy inference and triggers collection when their discrepancy indicates an out-of-distribution region. In a separate feedback loop, low online advantage predictions identify essential segments for refining a progress-based advantage estimator. The updated estimator then guides advantage-conditioned behavioral cloning using a balanced mixture of base demonstrations and new policy data. This design preserves the iterative and policy-aware nature of human-in-the-loop learning while decoupling data collection from robot deployment. Experiments on four real-world tasks spanning long-horizon and precise manipulation show that HIL-UMI achieves consistent improvement over SFT and benefits from both targeted collection and advantage refinement. Moreover, HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection time, suggesting a scalable path for VLA post-training across operators and locations.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20659v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20659v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Ownership in AI-Assisted Everyday Tasks</title>
      <link>https://arxiv.org/abs/2609.20658v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20658v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:38:35 GMT</pubDate>
      <dc:creator>Megan Wei, Melanie Subbiah, Audrey Lee et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>When does work done with AI still feel like ours? As AI becomes woven into everyday tasks, we must examine what happens to our sense of ownership and contribution when a machine shares in producing wh...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Ownership in AI-Assisted Everyday Tasks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Megan Wei, Melanie Subbiah, Audrey Lee et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">When does work done with AI still feel like ours? As AI becomes woven into everyday tasks, we must examine what happens to our sense of ownership and contribution when a machine shares in producing what we make. We report an exploratory qualitative survey in which participants were asked to describe two recent, self-selected tasks completed with AI: one that felt like their own and one that did not. We find that felt ownership depends on the process of collaboration: people disown work when they merely approve AI&apos;s suggestions, but retain ownership when they lead, iterate, or rewrite. Ownership can also extend to settings where people own the vision for a project but not the execution; respondents reported high ownership on tasks they could not have completed without AI. Loss of personal voice and a lack of comprehension of the output both erode ownership. Finally, willingness to disclose AI use is often decoupled from actual pride or ownership, and instead shaped by community norms and fear of credit erasure. We propose several research directions as a result of these findings to promote AI development that supports people&apos;s sense of authorship over their own lives.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20658v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20658v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning</title>
      <link>https://arxiv.org/abs/2609.20650v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20650v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:28:45 GMT</pubDate>
      <dc:creator>Simon Süwer, Julian Klemm, Elisa Acitelli et al.</dc:creator>
      <category>模型架构</category>
      <description>Federated learning enables collaborative training without sharing patient-level data, but most studies remain simulations. Based on five requirements derived from the literature, we analyzed 14 FL fra...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Simon Süwer, Julian Klemm, Elisa Acitelli et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Federated learning enables collaborative training without sharing patient-level data, but most studies remain simulations. Based on five requirements derived from the literature, we analyzed 14 FL frameworks and found that none fully satisfied these requirements. We present FL-Net, a novel federated clinical research framework to fulfill all requirements. It integrates modular data harmonization, data discovery, disclosure control, securely built versioned FL-Net-Tools and containerized federated workflow execution into a persistent network. It enables the re-use of harmonized data and workflows across studies. FL-Net&apos;s end-to-end capabilities were evaluated through harmonization, cross-study patient discovery across MIMIC and US-130, and reproducible, audited federated workflows with up to 50 concurrent clients. FL-Net is being developed within the dAIbetes and Microb-AI-ome EU projects and will cover over 800,000 patients across 10 hospitals in 9 countries covering longitudinal and single point in time data, FL-Net provides a practical foundation for interoperable, reproducible, and privacy-preserving multicenter clinical research.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20650v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20650v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation</title>
      <link>https://arxiv.org/abs/2609.20649v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20649v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:28:34 GMT</pubDate>
      <dc:creator>Yan Qin, Yue Chen, Wenwei Lin et al.</dc:creator>
      <category>模型架构</category>
      <description>Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. W...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yan Qin, Yue Chen, Wenwei Lin et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20649v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20649v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[评估与优化] [模型架构] [扩散模型] [图像生成] Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network</title>
      <link>https://arxiv.org/abs/2609.20633v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20633v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:19:28 GMT</pubDate>
      <dc:creator>Yulong Chen, Ziqian Zhang, Haoyu Zhang et al.</dc:creator>
      <category>评估与优化</category>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-based editors rely on spatial controls whose inaccuracies can leave edits incomplete...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yulong Chen, Ziqian Zhang, Haoyu Zhang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 评估与优化, 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, clip score, image editing, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-based editors rely on spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. We introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on a Generative Refinement Network. Our key idea is to couple edit localization with content generation through the global refinement of binary image codes, allowing editing evidence to be reassessed as the image evolves. RefineEdit initializes an editing branch from an intermediate source state, reusing the emerging layout. We compare the probabilities assigned by the two branches to the same source-sampled bits, using their signed differences to select editable positions and bits. Selected bits follow editing refinement, while the remaining bits copy the evolving source state. To stabilize editing across refinement steps, adaptive spatial freezing limits unnecessary mask expansion, while finite bit locking keeps recently selected bits editable. The framework requires no additional training, external masks, or attention control. Across nine editing categories of PIE-Bench, RefineEdit achieves the best background-preservation scores in PSNR, LPIPS, MSE and SSIM, together with the highest whole-image and edited-region CLIP scores among the evaluated methods.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20633v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20633v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 评估与优化, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies</title>
      <link>https://arxiv.org/abs/2609.20620v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20620v1</guid>
      <pubDate>Thu, 17 Sep 2026 16:03:13 GMT</pubDate>
      <dc:creator>Khalid Halba, Kylie Cooper, James G. Bellingham</dc:creator>
      <category>模型架构</category>
      <description>Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Khalid Halba, Kylie Cooper, James G. Bellingham</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnostic and recovery planner when onboard anomaly detection identifies performance outside expected limits. Because language models are stochastic, rigorous evaluation requires ensemble testing rather than individual demonstrations. We present a closed-loop simulation architecture that couples real-time C vehicle software with a higher-level orchestration layer for physics-based fault injection, structured prompting, language-model interaction, mission file generation, validation, execution, and LLM-judge scoring. The framework, which we call SPAR (Simulation Platform for AUV Recovery), supports evaluation across fault realizations, prompt structures, reasoning models, and mission conditions. We vary these for a mass-shift fault over 480 SPAR trials, evaluating a frontier model and three off-the-shelf locally deployable LLMs. Model choice dominates diagnosis: the frontier model places the CG-shift mechanism in its top three hypotheses in 85-90% of trials, versus 60-78% for the best local model. Reasoning analysis indicates that local-model success is associated with following the complete diagnostic procedure, whereas weaker models often commit prematurely to elevator failure even though the actuator tracks its command. Diagnosis and operational decision performance do not appear to be coupled in this dataset. The contributions are an architecture extending unanticipated-fault recovery from detection to mitigation and an ensemble methodology for evaluating LLM-assisted mission management on low-power AUVs.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20620v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20620v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression</title>
      <link>https://arxiv.org/abs/2609.20598v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20598v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:47:46 GMT</pubDate>
      <dc:creator>Zewen Yang, Xiaobing Dai, Zhenxiao Yin et al.</dc:creator>
      <category>模型架构</category>
      <description>In this paper, we tackle the problem of jointly estimating the system states and partially unknown dynamics within distributed sensor-equipped networks, particularly in scenarios where only partial st...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zewen Yang, Xiaobing Dai, Zhenxiao Yin et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">In this paper, we tackle the problem of jointly estimating the system states and partially unknown dynamics within distributed sensor-equipped networks, particularly in scenarios where only partial state observations are available. To address this issue, we propose an observer-based dynamic cooperative learning framework incorporating online distributed Gaussian Process (GP) regression, which enables accurate estimation despite incomplete in measurements and deficient GP models. In addition, a novel data collection strategy is introduced, with theoretical conditions ensuring feasible data acquisition. Moreover, we also derive an error upper bound encompassing state estimation and model estimation, leveraging the deterministic error bounds of GPs. Empirical simulations demonstrate the superiority of our approach compared to existing distributed GP-based methods.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20598v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20598v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] CrystalMO-TuRBO: Multi-Objective Trust-Region Bayesian Optimization for High-precision Joint Crystal Structure Refinement</title>
      <link>https://arxiv.org/abs/2609.20592v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20592v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:45:06 GMT</pubDate>
      <dc:creator>Joseph Agada, Yishu Wang, Arpan Biswas</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Crystal structure refinement is a fundamental inverse problem in materials characterization, where structural parameters are optimized to reproduce experimental diffraction data. Conventional approach...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">CrystalMO-TuRBO: Multi-Objective Trust-Region Bayesian Optimization for High-precision Joint Crystal Structure Refinement</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Joseph Agada, Yishu Wang, Arpan Biswas</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Crystal structure refinement is a fundamental inverse problem in materials characterization, where structural parameters are optimized to reproduce experimental diffraction data. Conventional approaches, such as least-squares and likelihood-based optimization, rely on local search and often struggle with non-convex, noisy, and highly correlated parameter landscapes, particularly when integrating multiple diffraction modalities. Joint refinement of X-ray and neutron data is especially challenging due to their complementary but competing sensitivities, which are typically combined through scalarized objectives requiring manual weighting and leading to suboptimal solutions. We propose CrystalMO-TuRBO, a multi-objective trust region Bayesian optimization architecture for joint crystal structure refinement. The method models X-ray and neutron discrepancies as separate objectives and transforms the problem into a normalized maximization setting. A two-phase optimization strategy is introduced: Phase 1 performs global exploration using parallel trust-region Bayesian optimization across multiple scalarizations to identify promising regions of the parameter space, while Phase 2 conducts localized refinement within a shrinking region to achieve high-precision solutions. This design explicitly separates global search from fine-grained optimization, addressing the unique accuracy requirements of refinement tasks. We evaluate the proposed method on experimentally collected X-ray and neutron diffraction data from single-crystal Ho2Ti2O7. Results demonstrate improved convergence, robustness, and parameter precision compared to classical refinement methods and Bayesian optimization baselines on refinement of a single-crystal pyrochlore material system.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20592v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20592v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding</title>
      <link>https://arxiv.org/abs/2609.20586v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20586v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:39:25 GMT</pubDate>
      <dc:creator>Zhikun Zhou, Kunyu Peng, Runyi Yang et al.</dc:creator>
      <category>模型架构</category>
      <description>Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such gro...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhikun Zhou, Kunyu Peng, Runyi Yang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent&apos;s observations, cooperative settings require this ability to remain effective after independently reconstructed maps are aligned and fused. In this setting, the referred target or its contextual landmark may come from another agent&apos;s observations, while spatial relations must still be interpreted from the querying robot&apos;s viewpoint. We formulate this problem as cooperative referring Gaussian grounding over fused maps, which requires geometric alignability, instance-level semantic comparability, and view-conditioned relation reasoning. Existing language-aware Gaussian methods mainly focus on single-map querying, whereas Gaussian registration methods optimize geometric or photometric alignment without preserving language-grounding-oriented semantic compatibility. We propose CoRef-GS, a cooperative referring Gaussian splatting framework. CoRef-GS constructs local open-vocabulary instance-aware Gaussian maps, then aligns partially overlapping maps with a cross-agent alignment module by geometric and semantic consistency, and grounds queries using a view-conditioned mask relation graph. We further introduce CoQuad-Ref, a dual-quadruped benchmark spanning both real-world and simulated indoor scenes. Experiments show that, on simulated scenes, CoRef-GS reduces the rotation error from 2.58° after coarse initialization to 0.15° after refinement, and improves real-world referring mIoU over ReferSplat from 52.6% to 68.8%. The established benchmark and source code will be publicly released at https://github.com/ruojiruoli17/CoRef-GS.git.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20586v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20586v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Limits of Confidence in Diffusion</title>
      <link>https://arxiv.org/abs/2609.20581v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20581v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:36:53 GMT</pubDate>
      <dc:creator>Russ Webb, Amitis Shidani, Alice Bizeul et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which p...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Limits of Confidence in Diffusion</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Russ Webb, Amitis Shidani, Alice Bizeul et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest (pixels, phonemes, or words) there are inherent dependencies between tokens. We show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, that no product of per-position distributions can match a dependent group, and that per-position distributions do not determine whether a group is dependent: two joint distributions can have identical per-position marginals while differing in which combinations of values occur. On ScanAndAdd, a synthetic task whose joint distribution is available in closed form, we verify that every group of two or more undetermined positions a confidence ranking writes is dependent, and measure the generated distribution to be $29\times$ the sampling-noise floor total variation while per-sample metrics are $1.0$.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20581v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20581v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] NS3Learn: Transferring 5G NR Mode-2 Reception Realism from ns-3 to the Veins/SUMO Stack for Connected-Vehicle Safety Assessment</title>
      <link>https://arxiv.org/abs/2609.20578v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20578v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:35:42 GMT</pubDate>
      <dc:creator>Rasheed Bello, Arthur Mukwaya, Gurcan Comert et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Connected-vehicle safety evaluations rely on coupled traffic and network simulations, but standard channel models ignore radio resource competition in 5G NR sidelink Mode-2, reporting unrealistically ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">NS3Learn: Transferring 5G NR Mode-2 Reception Realism from ns-3 to the Veins/SUMO Stack for Connected-Vehicle Safety Assessment</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Rasheed Bello, Arthur Mukwaya, Gurcan Comert et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Connected-vehicle safety evaluations rely on coupled traffic and network simulations, but standard channel models ignore radio resource competition in 5G NR sidelink Mode-2, reporting unrealistically high message delivery in dense traffic. This study introduces resource-competition losses without requiring full protocol reimplementation. We labeled 10.5 million reception outcomes from ns-3 5G-LENA traces (calibrated on 3GPP scenarios and driven by SUMO trajectories) to fit NS3Learn - a closed-form model capturing half-duplex loss, scheduling collisions, receiver capture, and decoding. Evaluation spanned two signalized urban networks, six penetration levels (1-100%), and five random seeds per condition. NS3Learn achieved a mean absolute deviation of 0.06 in per-instant delivery compared to ns-3 5G-LENA, outperforming alternative models (0.44 and 0.55 deviation). Fitted parameters transferred to a distinct intersection with only 20% additional error. Crucially, using realistic communication models reversed simulated traffic speed trends and more than doubled predicted hard-braking events. The framework transfers reception realism between simulators via model distillation instead of full reimplementation. Every stage maps directly to an explicit physical mechanism. Researchers and transportation agencies can maintain existing simulation pipelines while accurately accounting for dense-traffic packet loss and denial-of-service impacts. Adapting to new radio configurations requires only offline refitting rather than code modification.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20578v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20578v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] TAP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models</title>
      <link>https://arxiv.org/abs/2609.20577v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20577v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:34:40 GMT</pubDate>
      <dc:creator>Jingbo Liu, Zhiyuan Yu</dc:creator>
      <category>模型架构</category>
      <description>We study the Bayes-optimal spherical linear model as the ambient dimension and sample size grow proportionally, under a quantitative Marchenko--Pastur spectral-regularity condition on the design. This...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">TAP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jingbo Liu, Zhiyuan Yu</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We study the Bayes-optimal spherical linear model as the ambient dimension and sample size grow proportionally, under a quantitative Marchenko--Pastur spectral-regularity condition on the design. This condition is satisfied by normalized i.i.d. designs with standardized entries of finite fourth moment, but does not require entrywise independence or impose conditions on the singular vectors. Under this condition, we prove a quantitative all-temperature TAP approximation and characterize the posterior geometry. For the natural finite-aspect-ratio TAP functional, the normalized spherical free energy and the TAP optimum differ by $O_P(p^{-1})$. Each is within $O_P(p^{-1/2})$ of its explicit deterministic equivalent, and this fluctuation scale is sharp. Uniformly over all global TAP maximizers, the normalized squared Euclidean distance to the spherical posterior mean is $O_P(p^{-1})$. We also prove that the posterior mass outside a data-dependent band determined by the ridge estimator has sharp exponential order. More precisely, uniformly over sufficiently small band widths $\varepsilon$, the logarithm of this mass is at most $-cp\varepsilon^2+O_P(1)$. For every fixed geometrically admissible width, a spherical-cap construction gives a matching exponential-order lower bound on this mass. For every deterministic sequence of widths $\varepsilon_p\gg p^{-1/2}$, the corresponding bands capture asymptotically all posterior mass.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20577v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20577v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A Dual-Stream Regulated Reconstruction and Segmentation Network with Hierarchical Artifact-Prior Modeling for Ultra-Low-Field Pediatric Neuroimaging</title>
      <link>https://arxiv.org/abs/2609.20562v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20562v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:26:57 GMT</pubDate>
      <dc:creator>Bahram Jafrasteh, Leo Milecki, Qingyu Zhao</dc:creator>
      <category>模型架构</category>
      <description>Automated quality assessment, enhancement, and segmentation of multiple structures in $0.064\,\mathrm{T}$ ultra-low-field pediatric MRI are limited by a low signal-to-noise ratio, weak anatomical boun...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Dual-Stream Regulated Reconstruction and Segmentation Network with Hierarchical Artifact-Prior Modeling for Ultra-Low-Field Pediatric Neuroimaging</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Bahram Jafrasteh, Leo Milecki, Qingyu Zhao</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> u-net, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Automated quality assessment, enhancement, and segmentation of multiple structures in $0.064\,\mathrm{T}$ ultra-low-field pediatric MRI are limited by a low signal-to-noise ratio, weak anatomical boundaries, and frequent artifacts. We present a unified framework for the LISA 2026 Challenge that performs all three tasks together within one inference pipeline. A network with two coupled streams, built on a 3D U-Net, first reconstructs an enhanced uLF volume and then combines the original and enhanced images for subcortical segmentation. To improve boundary stability, we add an auxiliary class covering brain tissue outside the target structures, derived from whole brain masks. A head conditioned on an artifact graph predicts the seven artifact ratings from reconstruction residuals and frozen segmentation features. We address the scarcity of dense annotations using diffeomorphic registration from atlas to target for label propagation and to regularize anatomical reconstruction. We report validation results across all three tasks.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20562v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20562v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Mitigating Retaliatory Algorithmic Collusion in Repeated Games</title>
      <link>https://arxiv.org/abs/2609.20548v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20548v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:12:59 GMT</pubDate>
      <dc:creator>Karthik Sivachandran, Rohan Paleja</dc:creator>
      <category>模型架构</category>
      <description>Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared de...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Mitigating Retaliatory Algorithmic Collusion in Repeated Games</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Karthik Sivachandran, Rohan Paleja</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general repeated games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces a quantifiable conditional dependence in agents&apos; policies, detectable via the total variation distance between an agent&apos;s action distributions across cooperation and defection histories. Building on this connection, we propose CURB (Collusion Unwinding via Reward shaping and Belief injection), a reward-shaping framework that penalizes this Total Variation (TV) distance signal during Q-learning and is guaranteed to convert any SPC fixed point of the dynamics into a trivial one, thus precluding collusive equilibria sustained by punishment threats. Empirically, CURB substantially reduces collusion by Q-learning agents in both Bertrand and Cournot Competition Repeated Games. We further demonstrate that CURB extends to deep Q-network agents in Bertrand competition, suggesting the mechanism generalizes beyond tabular Q-learning.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20548v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20548v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Language-model groups overstate consensus when replaying human deliberation on a reasoning task</title>
      <link>https://arxiv.org/abs/2609.20543v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20543v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:11:08 GMT</pubDate>
      <dc:creator>Tengfei Shao</dc:creator>
      <category>模型架构</category>
      <description>Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with mat...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Language-model groups overstate consensus when replaying human deliberation on a reasoning task</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tengfei Shao</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant&apos;s pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20543v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20543v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Parallelism, critical windows, and separations among diffusion language models</title>
      <link>https://arxiv.org/abs/2609.20539v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20539v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:10:18 GMT</pubDate>
      <dc:creator>Sitan Chen, Liye Wang</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which r...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Parallelism, critical windows, and separations among diffusion language models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sitan Chen, Liye Wang</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, autoregressive model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among the many competing paradigms for dLLMs, from masked to uniform to Gaussian diffusion, principled understanding of how these different proposals compare in parallelism remains limited. In this work, we initiate a fine-grained comparison of the capacity for parallelism among these three leading approaches and prove the following:   - Uniform and Gaussian diffusion can sample in a number of forward passes which scales with the dual total correlation of the underlying distribution, a measure of intrinsic complexity which can be much smaller than the context length. Previously, it was only known how to achieve this using masked diffusion.   - For a certain family of random empirical measures, we show that $\widetildeΘ(\sqrt{d})$ forward passes are necessary and sufficient to sample using uniform or Gaussian diffusion, yet there exist approximate score oracles for which $\widetildeΩ(d)$ forward passes are needed for masked diffusion. This establishes the first provable separation in parallelism between the three prevailing dLLM paradigms.   Contrary to popular intuition that masked diffusions are harder to parallelize because they must commit to token values, the latter separation instead comes from the fact that the critical windows in masked diffusion sampling are asymptotically narrower than those in uniform and Gaussian diffusion sampling.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20539v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20539v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] FreqCondNorm: Towards Cross-domain Predictive Maintenance through a Frequency-Conditioned Transformer Foundation Model</title>
      <link>https://arxiv.org/abs/2609.20535v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20535v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:07:43 GMT</pubDate>
      <dc:creator>Zaynab Raounak, Camille LHermine, Zhiguo Zeng</dc:creator>
      <category>模型架构</category>
      <description>Deep learning predictive maintenance models suffer from poor transferability across machines and operating conditions, especially when labelled data are scarce and signals span five orders of magnitud...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">FreqCondNorm: Towards Cross-domain Predictive Maintenance through a Frequency-Conditioned Transformer Foundation Model</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zaynab Raounak, Camille LHermine, Zhiguo Zeng</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Deep learning predictive maintenance models suffer from poor transferability across machines and operating conditions, especially when labelled data are scarce and signals span five orders of magnitude in sampling frequency (1 Hz to ~100 kHz). We propose FreqCondNorm, a Transformer-based architecture that introduces a FiLM-style frequency-conditioned normalization layer to unify heterogeneous time-series within a single model. The architecture is pretrained on five public predictive maintenance datasets (CWRU, MFPT, UOC18, PRONOSTIA, CMAPSS) using masked auto-encoding and contrastive learning with balanced domain sampling. On fault diagnosis, the model achieves 99.2% accuracy on CWRU (+6.4 pp over CNN) and 82.1% zero-shot accuracy on MFPT, demonstrating strong transfer across sampling frequencies. However, the approach does not improve remaining useful life prediction, suggesting a mismatch between pretraining and RUL objectives that warrants future investigation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20535v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20535v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Relational Attention for Data-Efficient Language Modeling</title>
      <link>https://arxiv.org/abs/2609.20530v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20530v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:05:19 GMT</pubDate>
      <dc:creator>Adrian Brasoveanu, Ece Takmaz, Jakub Dotlačil</dc:creator>
      <category>模型架构</category>
      <description>We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replac...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Relational Attention for Data-Efficient Language Modeling</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Adrian Brasoveanu, Ece Takmaz, Jakub Dotlačil</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a Dual Attention Transformer (DAT), which separates the routing of object-level (&quot;sensory&quot;) lexical features from structural/relational information (Altabaa and Lafferty, 2025; Altabaa et al., 2024; Webb et al., 2024; Kerg et al., 2022; Webb et al., 2021). Relational attention (RA) disentangled from self-attention greatly increases data efficiency and out-of-training-sample generalization on purely relational tasks, but language modeling requires object-level and relational information to be integrated as well as disentangled, and RA-based LMs have remained largely unexplored. BabyLM&apos;s data-constrained training and comprehensive evaluation is an ideal testing ground for whether that data efficiency transfers. As a training intervention, we add a Next-Latent Prediction (NextLat; Teoh et al. 2026) objective that encourages hidden states to compress history incrementally into a dense belief state. Architecture is the dominant factor for structural linguistic generalization; the objective is secondary but still significant. DAT&apos;s three relational attention types (full RA vs. the simpler RCA and DisRCA variants) are largely interchangeable at 10M words; full RA pulls ahead at 100M. We also introduce a novel symbol-retrieval mechanism (RoPE-based, as opposed to learned, relative symbols) that matches learned symbol libraries while adding no parameters. On the strict (100M-word) track, our best model ranks 6th of 55 overall and 3rd of 55 on the leaderboard&apos;s NLP-task subset at the time of writing; our two strongest models outperform the GPT-2 baseline on most benchmarks, with one attaining the highest EWoK score among strict-track entries.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20530v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20530v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning</title>
      <link>https://arxiv.org/abs/2609.20523v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20523v1</guid>
      <pubDate>Thu, 17 Sep 2026 15:01:02 GMT</pubDate>
      <dc:creator>Bo Tang, Zixuan Liao, Hao Li et al.</dc:creator>
      <category>模型架构</category>
      <description>Quantum communication underpins secure information processing and scalable quantum networks. In particular, remote state preparation (RSP) enables efficient quantum state transfer, but accurately esti...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Bo Tang, Zixuan Liao, Hao Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Quantum communication underpins secure information processing and scalable quantum networks. In particular, remote state preparation (RSP) enables efficient quantum state transfer, but accurately estimating target states under complex noise remains challenging. Here, we propose a Transformer-based Quantum State Characterizer (TQSC) model for noisy RSP experiments. Our model reconstructs experimentally prepared pure and mixed photonic polarization states from noisy measurements in complex scattering environments, while its attention patterns provide physically grounded insights into correlations among the measured observables. The method achieves a mean estimator-target fidelity exceeding 99.999% under complex scattering and dynamic Gaussian noise, while its robustness and generalization are further examined using Qiskit-simulated Bloch-ball states.Furthermore, in a practical MNIST image transmission task with held-out states, the decoded bit error rate is reduced from 50.34% to zero after TQSC post-processing. The TQSC model enables accurate tomographic characterization under dynamic noise and provides physically grounded post-hoc insights, holding promise for intelligent quantum information processing applications.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20523v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20523v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness</title>
      <link>https://arxiv.org/abs/2609.20519v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20519v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:58:29 GMT</pubDate>
      <dc:creator>Haozhe Liu, Tian Ye, Sensen Gao et al.</dc:creator>
      <category>图像生成</category>
      <description>As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedb...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Haozhe Liu, Tian Ye, Sensen Gao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvements that transfer beyond their development setting, moving automated harness discovery toward production-level outcomes. Four mechanisms survive selection and form SoL-Pi, spanning action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, SoL-Pi achieves performance comparable to Pi across GPT-5.6 Sol and Opus 5 while reducing recorded token traffic by 44.7-49.0% and API cost by about one third. In other words, estimated hourly savings are \$8.75-\$13.50 relative to native Codex and Claude Code harnesses, and \$4.36-\$5.71 relative to Pi.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20519v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20519v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation</title>
      <link>https://arxiv.org/abs/2609.20511v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20511v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:53:39 GMT</pubDate>
      <dc:creator>Yuxiao Yang, Tianrun Yu, Shangzhe Li et al.</dc:creator>
      <category>扩散模型</category>
      <description>We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination-token mismatch} bet...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yuxiao Yang, Tianrun Yu, Shangzhe Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination-token mismatch} between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student&apos;s preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning the decoding stopping set alone is insufficient, while treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced length inflation across all three model families. To further understand how termination behavior evolves over training, we study OPD across different K2-Horizon training stages. This stage-wise analysis shows that termination preferences can shift substantially during training, while also revealing a distinct length inflation late in the OPD run that persists beyond termination alignment. Together, these results identify termination mismatch as an important, but not exhaustive, source of OPD length dynamics. We release an implementation incorporating the proposed termination-handling corrections.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20511v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20511v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Automated Goldsmith&apos;s Mark Retrieval in Silverware</title>
      <link>https://arxiv.org/abs/2609.20509v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20509v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:51:46 GMT</pubDate>
      <dc:creator>Atmik Tiwari, Vincent Christlein, Mark Fichtner et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>For art historians, goldsmith marks play a critical role in the identification and dating of artifacts. In practice, experts must manually compare a query mark against hundreds of documented examples,...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Automated Goldsmith&apos;s Mark Retrieval in Silverware</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Atmik Tiwari, Vincent Christlein, Mark Fichtner et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, imagen</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">For art historians, goldsmith marks play a critical role in the identification and dating of artifacts. In practice, experts must manually compare a query mark against hundreds of documented examples, a process that is both tedious and highly dependent on specialist knowledge. To address this, we present an AI-assisted retrieval pipeline that combines mark localization with metric-learning fine-tuning across three backbone architectures: an ImageNet-pretrained ResNet-50, a supervised ViT-S/16, and a self-supervised DINOv2 ViT-S/14. We conduct a systematic evaluation of cropping strategies, where we measure the impact of no cropping, manual ground-truth cropping, and learned detection-based cropping, and assess their interaction with each backbone. Our strongest configuration, DINOv2 ViT-S/14 with manual crop and metric-learning fine-tuning, achieves an mAP of 62.63% and a Top-1 accuracy of 73.74%. Our experiments show that self-supervised pretraining and mark localization are the two most impactful factors, with learned cropping recovering the majority of the gain from manual cropping without requiring ground-truth annotations at inference time. To enable reproducibility and adoption in the digital humanities, we release our manually annotated dataset and codebase, and deploy the system via a public web interface.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20509v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20509v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain</title>
      <link>https://arxiv.org/abs/2609.20504v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20504v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:50:12 GMT</pubDate>
      <dc:creator>Aakash Singh, Lakshmi Pedapudi, Chandrashekar M S et al.</dc:creator>
      <category>模型架构</category>
      <description>FarmerChat is Digital Green&apos;s AI-powered agricultural advisory assistant for smallholder farmers, who access it in their own language through text, voice, or photographs. Voice is a critical channel f...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Aakash Singh, Lakshmi Pedapudi, Chandrashekar M S et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">FarmerChat is Digital Green&apos;s AI-powered agricultural advisory assistant for smallholder farmers, who access it in their own language through text, voice, or photographs. Voice is a critical channel for this population, yet field-recorded speech is challenging for general-purpose automatic speech recognition (ASR) because recordings frequently contain machinery noise, background media, competing speakers, and domain-specific agricultural vocabulary. These conditions disproportionately affect crop, pest, chemical, and quantity terms that carry the meaning of a farmer&apos;s query.   We present a modular, model-agnostic pipeline for improving ASR quality in FarmerChat without fine-tuning or replacing the underlying ASR model. The pipeline combines gated audio enhancement, speaker diarization and target-speaker selection, ASR, domain-aware correction using a weighted agricultural lexicon, and a quality gate for detecting unreliable transcripts. Only the diarization stage is fine-tuned; all other stages use off-the-shelf models behind common interfaces.   We evaluate the pipeline on human-annotated FarmerChat recordings in Hindi, Telugu, and Odia using word error rate (WER) and a domain-weighted error rate that gives greater importance to agricultural terminology. The largest improvements occur on multi-speaker recordings, where target-speaker selection prevents competing speech from entering the transcript. Across the full corpus, the pipeline reduces WER by 16-23% relative on three cloud ASR models and by 5% on an on-device model. On multi-speaker recordings, the reductions are 32-42% for the cloud models and 16% for the on-device model. All reported reductions are statistically significant. These results show that targeted preprocessing, speaker selection, and domain-aware post-processing can substantially improve agricultural speech transcription while preserving the underlying ASR model.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20504v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20504v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Distributionally Robust Federated Learning with Multi-Source Data</title>
      <link>https://arxiv.org/abs/2609.20501v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20501v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:48:22 GMT</pubDate>
      <dc:creator>Yingzhu Liu, Zhongkui Li, Pengcheng You et al.</dc:creator>
      <category>模型架构</category>
      <description>Federated learning trains a shared model from private client data. In practice, data-generating distributions may differ, and the true mixture across clients is often unknown, making the underlying gr...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Distributionally Robust Federated Learning with Multi-Source Data</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yingzhu Liu, Zhongkui Li, Pengcheng You et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Federated learning trains a shared model from private client data. In practice, data-generating distributions may differ, and the true mixture across clients is often unknown, making the underlying group distribution difficult to specify. Existing approaches address cross-client mixture uncertainty by optimizing against the worst-case mixture, yet assume accurate client-wise distribution estimates. However, these estimates can be unreliable when based on finite samples. To handle both cross-client mixture uncertainty and within-client distributional ambiguity, we construct a global ambiguity set as the union of admissible mixtures of local ambiguity sets. The construction allows client-specific ambiguity radii and admits a client-wise separable reformulation. Leveraging this structure, we establish a high-probability out-of-sample performance guarantee. We further develop a federated algorithm for a penalty-based reformulation and prove its convergence under milder regularity conditions. Simulations validate the algorithm&apos;s effectiveness.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20501v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20501v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Resolution limits for process comparison from event data</title>
      <link>https://arxiv.org/abs/2609.20489v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20489v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:39:09 GMT</pubDate>
      <dc:creator>Antony R. Lee, Peter Tiňo, Iain B. Styles</dc:creator>
      <category>模型架构</category>
      <description>One hospital runs bloods and imaging at the same time. Another runs them one after the other, in either order, equally often. Knowing which actually happened, and how it is recorded in data, is critic...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Resolution limits for process comparison from event data</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Antony R. Lee, Peter Tiňo, Iain B. Styles</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">One hospital runs bloods and imaging at the same time. Another runs them one after the other, in either order, equally often. Knowing which actually happened, and how it is recorded in data, is critical for all operational managers. In process mining, the standard approach is to construct an event log, and attempt to discover concurrent and sequential processes in a data-driven way. We show this standard approach, built on the stochastic language of an event log, reports only the assumptions of its discovery algorithm, because every such log is explained equally well by a model with no concurrency at all. Further, before any data is acquired, we characterise when data can and cannot distinguish concurrent behaviour. Where it cannot, the distinction is recoverable from evidence the stochastic language discards, such as the times at which activities start and end, or object-centric records that fix an order within an execution. The remedy is therefore a choice of what is recorded, rather than a larger sample. This impacts decision making, as planning resource for truly concurrent services is very different from sequential services.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20489v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20489v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI</title>
      <link>https://arxiv.org/abs/2609.20481v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20481v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:34:34 GMT</pubDate>
      <dc:creator>Justin Payan, Bálint Gyevnár, Atoosa Kasirzadeh et al.</dc:creator>
      <category>模型架构</category>
      <description>Conferences, journals, funders, schools, and universities are struggling with a surge of potentially AI-generated submissions from ostensibly human authors, who may not have exercised sufficient human...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Justin Payan, Bálint Gyevnár, Atoosa Kasirzadeh et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Conferences, journals, funders, schools, and universities are struggling with a surge of potentially AI-generated submissions from ostensibly human authors, who may not have exercised sufficient human oversight for their manuscripts. In turn, institutions evaluating submissions can no longer reliably credit expertise based solely on authors&apos; names on submitted work. To address this problem, we propose greCAPTCHA, a proctored assessment approach that measures authors&apos; understanding of research manuscripts via the construct of capacity to verify, which we define as the knowledge and reasoning required to critically assess the contents underlying one&apos;s contributions to a manuscript. greCAPTCHA generates questions assessing multiple levels of understanding and provides an evaluative report based on authors&apos; responses. Using a prototype implementation, we conduct a user study and semi-structured interviews with $31$ researchers to evaluate greCAPTCHA. Its automated scores predict which papers were or were not authored by study participants with an AUC of $0.90$. Participants reported positive overall experiences with the system and remarked on the appropriate construct validity for author understanding, while also suggesting important changes to be made before deployment. Our results provide initial evidence that greCAPTCHA can assess manuscript-specific understanding under proctored conditions.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20481v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20481v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation</title>
      <link>https://arxiv.org/abs/2609.20475v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20475v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:32:00 GMT</pubDate>
      <dc:creator>Euiseok Han, Tri Ton, Hwanhee Kim et al.</dc:creator>
      <category>模型架构</category>
      <description>Open-vocabulary scene understanding is fundamental for robotics, laying the groundwork for spatial reasoning and object manipulation. While closed-vocabulary 3D instance segmentation heavily leverages...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Euiseok Han, Tri Ton, Hwanhee Kim et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Open-vocabulary scene understanding is fundamental for robotics, laying the groundwork for spatial reasoning and object manipulation. While closed-vocabulary 3D instance segmentation heavily leverages 3D shape information, state-of-the-art open-vocabulary methods remain predominantly restricted to 2D image features or image-distilled representations during mask labeling. In this paper, we propose SenseFuse, a label-free fusion method that balances 2D image and 3D shape encoders for robust open-vocabulary 3D instance segmentation, refining only the mask-labeling stage of existing pipelines. We reveal that 2D image and 3D shape encoders exhibit largely disjoint failure patterns and rarely share identical wrong labels, whereas two 2D image encoders frequently repeat the same errors. This distinct behavior makes the 2D and 3D pair inherently complementary. We introduce an adaptive mechanism that selects a scene-level fusion weight to maximize a label-free sensitivity measure, estimated directly from a single scene&apos;s unlabeled proposals in milliseconds. SenseFuse improves labeling accuracy in every evaluated setting across ScanNet200, Replica, and ScanNet++, recovering 67-100% (median 93%) of the gain achievable with an oracle weight, and it raises instance AP in 21 of 22 reported settings. Code is available at https://github.com/hanes1207/SenseFuse.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20475v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20475v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents</title>
      <link>https://arxiv.org/abs/2609.20474v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20474v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:30:14 GMT</pubDate>
      <dc:creator>Yukun Zhang, Kemu Xu, Yishen Chen</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airlin...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yukun Zhang, Kemu Xu, Yishen Chen</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier&apos;s avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20474v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20474v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Deep Learning-Based Classification of Cognitive and Resting States Using Electroencephalography Signals</title>
      <link>https://arxiv.org/abs/2609.20467v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20467v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:26:12 GMT</pubDate>
      <dc:creator>K. A. Januka S. Fernando, Harshit Srivastava</dc:creator>
      <category>模型架构</category>
      <description>The categorization of cognitive and resting states derived from electroencephalography (EEG) signals is crucial for comprehending fluctuations in brain activity linked to various mental states. EEG pr...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Deep Learning-Based Classification of Cognitive and Resting States Using Electroencephalography Signals</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> K. A. Januka S. Fernando, Harshit Srivastava</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The categorization of cognitive and resting states derived from electroencephalography (EEG) signals is crucial for comprehending fluctuations in brain activity linked to various mental states. EEG provides a non-intrusive approach for documenting brain function in both resting and task-oriented cognitive conditions, whilst deep learning techniques enable the automatic extraction of significant patterns from intricate EEG data. This study presents a deep learning framework to distinguish between resting and cognitive states through EEG records. The proposed framework integrates a Convolutional Neural Network (CNN) stacked with a Gated Recurrent Unit (GRU) for the extraction of features from EEG signals. Time-frequency analysis is conducted to explore the salient aspects of signals, and the derived features are then assessed utilizing conventional deep learning and machine learning classifiers, including the suggested 2D-Net architecture. The proposed approach and feature extraction strategy outperform the evaluated comparative methods, achieving accuracies of 83.177% for resting-versus-mathematical task classification, 76.107% for resting-versus-memory task classification, and 83.432% for resting-versus-music task classification. The findings illustrate the efficacy of integrating signal processing with deep learning methodologies to discriminate resting from cognitive states utilizing EEG signals.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20467v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20467v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Correlation-Free Transition Path Sampling through Shooting Point Generation Guided by Committor Learning</title>
      <link>https://arxiv.org/abs/2609.20461v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20461v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:24:14 GMT</pubDate>
      <dc:creator>Maximilian Negedly, Sebastian Falkner, Alessandro Coretti et al.</dc:creator>
      <category>模型架构</category>
      <description>Studying the dynamical behavior of a system often depends on characterizing how it transitions between long-lived states. Because such transitions are rare, observing them usually requires specialized...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Correlation-Free Transition Path Sampling through Shooting Point Generation Guided by Committor Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Maximilian Negedly, Sebastian Falkner, Alessandro Coretti et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Studying the dynamical behavior of a system often depends on characterizing how it transitions between long-lived states. Because such transitions are rare, observing them usually requires specialized enhanced sampling techniques. Transition Path Sampling (TPS) is a well-established method for generating reactive trajectories, which is simple to implement and does not require the definition of a preconceived reaction coordinate. However, its efficiency is limited by its sequential nature and the resulting correlations between sampled paths. Previous work addressed this limitation by combining TPS with a sampling scheme based on conditioned Boltzmann Generators, a generative machine learning model capable of sampling a given target probability distribution. This approach produces uncorrelated transition paths but relies on an accurate reaction coordinate, which is rarely known in advance. Building on recent advances in committor learning, specifically on the Artificial Intelligence for Molecular Mechanism Discovery (AIMMD) method, in this work we introduce GenAIMMD, an iterative algorithm that actively and self-consistently learns the ideal reaction coordinate (the committor) and trains a conditioned Boltzmann Generator to sample from arbitrary bias windows along it. GenAIMMD thereby provides a correlation-free and fully parallelizable path sampling scheme that does not require prior knowledge of the system&apos;s transition mechanism. We apply GenAIMMD to a two-dimensional toy model and a higher-dimensional polymer system. In both cases, GenAIMMD succeeds in training the Boltzmann Generator and learning the committor. Benchmark results show a substantial increase in performance compared to standard TPS.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20461v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20461v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Fingerprinting Multimodal Large Language Models</title>
      <link>https://arxiv.org/abs/2609.20457v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20457v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:21:39 GMT</pubDate>
      <dc:creator>Chao Huang, Meng Tong, Kejiang Chen</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Fingerprinting Multimodal Large Language Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Chao Huang, Meng Tong, Kejiang Chen</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in MLLMs and struggle to detect violations of distillation. To bridge this gap and safeguard model ownership, we present the first study on multimodal model fingerprinting. Inspired by recent findings that self-attention acts as a low-pass filter and that its low-frequency components are informative, we develop AttnPrint for white-box provenance. Specifically, we extract cross-modal attention distributions and isolate their low-frequency components to serve as model fingerprints. To facilitate black-box auditing, we further introduce DistillTrace, which employs hypothesis testing of MLLM outputs to identify potential model infringement. We conduct extensive experiments on 154 model instances across 19 multimodal architectures. Notably, AttnPrint achieves strong derivative-model detection performance while remaining robust to five downstream modification techniques. DistillTrace also provides evidence of distillation relationships under three parameter-independent techniques.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20457v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20457v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback</title>
      <link>https://arxiv.org/abs/2609.20455v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20455v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:20:45 GMT</pubDate>
      <dc:creator>Ziqiao Shang, Ling-Yue Ge, Lan-Zhe Guo</dc:creator>
      <category>模型架构</category>
      <description>External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an edit...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ziqiao Shang, Ling-Yue Ge, Lan-Zhe Guo</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation. We introduce SkillAA (Skill Abductive Attribution), a structured skill-optimization framework for frozen language models. It represents skill applicability, execution, and composition in a unified graph, allowing the same structure to support skill selection, attribution-guided repair, and update validation. SkillAA contrasts successful and failed executions to route candidate repairs to specific graph objects, updates only the selected local structure, and uses Local and Big Gates to screen candidate changes before commitment. With gpt-5.6-sol, SkillAA reaches 81.5%, 66.7%, and 91.2% on SearchQA, LiveMath, and DocVQA, respectively, and attains the highest observed mean in every main setting. These results support the utility of attribution-guided graph editing and graph-scoped validation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20455v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20455v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] The Organization of Inference: Information, Resource Constraints, and AI Production</title>
      <link>https://arxiv.org/abs/2609.20449v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20449v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:17:28 GMT</pubDate>
      <dc:creator>Yukun Zhang, Kemu Xu, Yishen Chen</dc:creator>
      <category>图像生成</category>
      <description>The economic value of inference depends on how capacity and task information are distributed across stages of AI production. We study these organizational margins using controlled workflow experiments...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">The Organization of Inference: Information, Resource Constraints, and AI Production</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yukun Zhang, Kemu Xu, Yishen Chen</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The economic value of inference depends on how capacity and task information are distributed across stages of AI production. We study these organizational margins using controlled workflow experiments on externally verified software-engineering tasks. In two matched resource panels, direct execution records the same success rate of 59.6 percent at logical-token ceilings of 12,000 and 24,000, while success under information-constrained planning rises from 36.2 to 51.2 percent. The planning disadvantage narrows by 15.0 percentage points (95 percent task-cluster bootstrap interval: 4.2 to 25.8). A strict read-only planning campaign varies whether the planner sees the task issue. At 12,000 tokens, issue access raises success by about 16 percentage points over issue-hidden planning. Compared with direct execution, task-informed planning is about 10 points lower at 12,000 tokens; at 24,000 tokens, it shows a 29.6-point advantage. In the resource panels, direct execution uses substantially less than either ceiling, while the planning workflow&apos;s binding rate falls from 46.2 to 0.8 percent and downstream execution accounts for 89.9 percent of the increase in total use. Scale determines the capacity available to a system; workflow and information structure shape the productive value</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20449v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20449v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Cross-Architecture Foundation-Model Distillation for Edge Flood Segmentation</title>
      <link>https://arxiv.org/abs/2609.20441v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20441v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:15:01 GMT</pubDate>
      <dc:creator>Fabian Schmalstieg, Karsten Mueller, Wojciech Samek</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Geospatial foundation models can provide strong flood-segmentation performance, but their size limits deployment on memory-constrained edge hardware. We distill a 300-million-parameter Prithvi-EO-2.0 ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Cross-Architecture Foundation-Model Distillation for Edge Flood Segmentation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Fabian Schmalstieg, Karsten Mueller, Wojciech Samek</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Geospatial foundation models can provide strong flood-segmentation performance, but their size limits deployment on memory-constrained edge hardware. We distill a 300-million-parameter Prithvi-EO-2.0 teacher, fine-tuned on the 252 manually labeled Sen1Floods11 training scenes, into a 0.7-million-parameter EfficientViT-B0 student. The teacher supervises additional unlabeled Sentinel-2 imagery, allowing the student training set to grow without new manual annotations. At the matched budget of 252 scenes, teacher-supervised training is competitive with direct training and improves STURM-Flood performance across tested configurations; a geometry-matched control shows that label source alone does not explain the difference. Scaling the teacher-supervised pool to 2,500 scenes narrows the remaining student--teacher gap: the float student reaches 0.787 water intersection over union on the Sen1Floods11 test split against 0.822 for the teacher, matches the teacher on STURM-Flood under our evaluation protocol, and remains below it on WorldFloods-v2. After activation replacement and quantization-aware training, the student runs as a 1.5-megabyte 8-bit integer (INT8) TensorRT engine on a Jetson Xavier NX at 5.57 milliseconds of graphics processing unit (GPU) compute per 512-by-512 image, with approximately 14 megabytes of runtime device memory. A fixed modified normalized difference water index (MNDWI) threshold is competitive with both models on the two clean external benchmarks, so we interpret those benchmarks as generalization tests rather than as evidence of learned-model superiority over a spectral rule. The results support the conclusion: foundation-model supervision can amplify a fixed manual annotation budget into a substantially larger training set and yield a compact, deployable edge model.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20441v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20441v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] A Mathematical Model of Motivated Emotional Mind - Cognitive Embodied System</title>
      <link>https://arxiv.org/abs/2609.20437v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20437v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:13:10 GMT</pubDate>
      <dc:creator>Wiesław L. Galus, Janusz A. Starzyk</dc:creator>
      <category>图像生成</category>
      <description>This article presents a mathematical model of the Motivated Emotional Mind cognitive architecture developed for embodied intelligent systems. Such a system learns to maintain its homeostasis through a...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Mathematical Model of Motivated Emotional Mind - Cognitive Embodied System</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Wiesław L. Galus, Janusz A. Starzyk</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">This article presents a mathematical model of the Motivated Emotional Mind cognitive architecture developed for embodied intelligent systems. Such a system learns to maintain its homeostasis through a generalized form of reinforcement learning based on its internal motivations, termed motivated learning (ML). The principal contribution of this article is a rigorous formalization of the re-entrant loop integrating feedforward processing, lateral interactions, and feedback pathways, together with the representational selection mechanisms that govern adaptive system responses. The model specifies how ongoing exteroceptive and interoceptive signals, bodily-motivational context, and memory traces are bound into associative memory structures termed semblions, which compete for access to further processing and top-down reconstruction. The formalization encompasses secondary perception, representational competition, curiosity, procedural gaps, and action selection directed toward limiting allostatic violations. Within this framework, motivated learning is tailored to embodied systems whose dynamics are shaped by needs, affect, and the current regulatory state. Unlike standard reinforcement-learning models, the proposed approach incorporates need thresholds, goal generation and shifting goals, bodily state, resource constraints, and action uncertainty, thereby providing a more adequate account of response selection under regulatory pressure. Global affect functions as a central control signal, modulating the learning rate, representational valence, and the balance between exploration and exploitation. The model presented here is a step toward a more rigorous formalization of cognitive phenomena and may provide a basis for further theoretical analysis, computer simulation, and implementation in artificial-intelligence systems inspired by biological processes.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20437v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20437v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain</title>
      <link>https://arxiv.org/abs/2609.20427v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20427v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:08:16 GMT</pubDate>
      <dc:creator>Alam Noor, Miguel Guti&apos;errez Gait&apos;an</dc:creator>
      <category>模型架构</category>
      <description>Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without j...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Alam Noor, Miguel Guti&apos;errez Gait&apos;an</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without justification. We ground a model in the Sheep Pain Facial Expression Scale (SPFES) by letting each detected facial region attend over text embeddings of the clinical descriptors and then test whether the resulting explanations mean anything. They do not. Ablating an entire descriptor changes the predicted logit by about $10^{-4}$, and the most-attended cue agrees with the predicted pain level in only $32.6\%$ of regions, although the attention maps, the learned gate, and the generated text all proposed otherwise. We therefore remove the appearance bypass with a concept bottleneck whose classifier reads only SPFES concept scores, supervised by per-region state annotations that image-level pipelines discard. This costs $0.05$--$0.10$ in Cohen&apos;s $κ$ but yields concepts that are demonstrably learned: minority pain-indicating states are recovered at $3.5$--$8.3\times$ their base rates, and the ear and eye severity orderings emerge without severity supervision. Removing the supervision alone leaves $κ$ unchanged while concept accuracy falls to $0.109$, showing that architectural necessity does not imply semantic validity. We also show that pooled concept accuracy is misleading under clinical imbalance and provide a cross-validated, protocol-matched benchmark of seven methods on this dataset.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20427v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20427v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing</title>
      <link>https://arxiv.org/abs/2609.20423v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20423v1</guid>
      <pubDate>Thu, 17 Sep 2026 14:04:31 GMT</pubDate>
      <dc:creator>Hao Yu, Kang Liu, Linnan Zhao et al.</dc:creator>
      <category>模型架构</category>
      <description>Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common doc...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Hao Yu, Kang Liu, Linnan Zhao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser&apos;s remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to measure the Stage I parser&apos;s residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20423v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20423v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] The Bias of Nonlinear Two-Time-scale Stochastic Approximation under Constant Step-Sizes</title>
      <link>https://arxiv.org/abs/2609.20409v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20409v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:54:20 GMT</pubDate>
      <dc:creator>Djamel Rassem Lamouri, Dorian Baudry, Nicolas Gast</dc:creator>
      <category>模型架构</category>
      <description>Two-timescale stochastic approximation (TTSA) is a fundamental tool for analyzing coupled iterative algorithms in reinforcement learning, optimization, and stochastic control. However, finite-time gua...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">The Bias of Nonlinear Two-Time-scale Stochastic Approximation under Constant Step-Sizes</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Djamel Rassem Lamouri, Dorian Baudry, Nicolas Gast</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Two-timescale stochastic approximation (TTSA) is a fundamental tool for analyzing coupled iterative algorithms in reinforcement learning, optimization, and stochastic control. However, finite-time guarantees for nonlinear two-timescale schemes remain difficult to obtain, especially under constant step-sizes. In this paper, we study nonlinear TTSA with step-sizes $α\ggβ$. Under standard stability, regularity, and Markovian noise assumptions, we upper bound the mean-squared error and the bias of both iterates around their limiting equilibria. Our bounds scale as $O(α+β^2/α^2)$, which we prove to be tight when $β\leα^{3/2}$. The analysis separates the contributions of initial conditions, fast-timescale tracking error, Markovian dependence, and timescale coupling, thereby clarifying the origin of the $β^2/α^2$ term. Our results reveal qualitative differences from the linear TTSA setting previously studied, showing that nonlinear dynamics introduce additional finite-time effects that are absent in the linear case.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20409v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20409v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection</title>
      <link>https://arxiv.org/abs/2609.20404v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20404v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:53:00 GMT</pubDate>
      <dc:creator>Rishi Bharadwaj, Yadati Narahari</dc:creator>
      <category>模型架构</category>
      <description>Agricultural soils are a major untapped carbon sink. Carbon farming is emerging as a promising practice for tapping this potential. Smallholder farmers, who dominate agriculture across South Asia and ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Rishi Bharadwaj, Yadati Narahari</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Agricultural soils are a major untapped carbon sink. Carbon farming is emerging as a promising practice for tapping this potential. Smallholder farmers, who dominate agriculture across South Asia and sub-Saharan Africa, are key to scaling climate mitigation via carbon farming. It is ironic that real-world carbon programs largely fail to reach them. We study this important gap through the lens of contract design. An aggregator offers a single pooled contract to a heterogeneous population of smallholder farmers who have private adoption costs (adverse selection) and exert unobserved effort (moral hazard), with agronomic outcomes evolving over multiple seasons. We formulate this evolving contracting problem as a POMDP and use reinforcement learning to learn a dynamic profit-maximising contract. We analyse the performance of the aggregator under various conditions. We find that a profit-maximising aggregator does not merely inherit the exclusion of smallholders, it amplifies it. On large farms the aggregator realises 87.7% of achievable adoption, against only 8.2% on smallholdings. Per-hectare Measurement, Reporting and Verification (MRV) costs fall as farm size rises, and the aggregator&apos;s pooling contract compounds this gradient rather than offsetting it. A counterfactual that makes MRV costs purely area-proportional eliminates this disparity. Our results and simulation can guide contract and policy design that opens carbon income to smallholders while enabling agricultural soils to contribute to climate mitigation at scale.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20404v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20404v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning</title>
      <link>https://arxiv.org/abs/2609.20389v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20389v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:41:29 GMT</pubDate>
      <dc:creator>Weiwei Wang, Yuqiang Li, Xianyi Wu et al.</dc:creator>
      <category>模型架构</category>
      <description>Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone a...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Weiwei Wang, Yuqiang Li, Xianyi Wu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone are insufficient; principled uncertainty quantification, such as confidence intervals and variance estimates, is essential for safe and risk-aware decision-making. A comprehensive way to unify these tasks is to estimate the sampling distribution of the evaluation error. Existing approaches, however, often suffer from limited robustness, scalability, or finite-sample validity. In this paper, we propose a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs). Unlike classical bootstrap methods that rely on resampling complete episodes, the proposed method regenerates trajectories from an estimated MDP and can therefore accommodate a much broader range of offline data formats, including complete trajectories, transition-level observations, and trajectory fragments. This flexibility further improves finite-sample statistical efficiency. We establish bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy value. Extensive simulations show that the proposed method accurately captures the sampling distribution of the OPE estimator, yielding tighter confidence intervals and more accurate variance estimates in most settings.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20389v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20389v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Navi-Agent: Unlocalized Monocular Navigation Agent</title>
      <link>https://arxiv.org/abs/2609.20388v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20388v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:41:27 GMT</pubDate>
      <dc:creator>Wenyuan Xie, Mengyang Hong, Yongzhong Wang et al.</dc:creator>
      <category>图像生成</category>
      <description>Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to execute long-horizon instructions in unknown environments. Existing zero-shot VLN-CE systems typically main...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Navi-Agent: Unlocalized Monocular Navigation Agent</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Wenyuan Xie, Mengyang Hong, Yongzhong Wang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to execute long-horizon instructions in unknown environments. Existing zero-shot VLN-CE systems typically maintain spatial states through geometric localization or coordinate-based representations. Recent geometry-constrained navigation removes depth and globally consistent coordinates, but maintaining persistent spatial awareness for place confirmation, progress verification, and recovery remains challenging. We present Navi-Agent, a zero-shot VLN-CE agent that constructs a coordinate-free spatial state from visual observations and executed motion histories. Navi-Agent organizes this state as a navigation topology, where nodes represent visual places and edges represent motion transitions. This representation enables observation-based approximate self-localization, task progress verification, and visual revisitation-based recovery. Navi-Agent performs closed-loop navigation by decomposing instructions into sub-goals, executing local visual navigation, and verifying visited places through the constructed spatial state. Experiments on zero-shot VLN-CE benchmark and real-world robot platforms show that Navi-Agent achieves state-of-the-art performance among geometry-constrained methods while remaining competitive with approaches relying on geometric localization.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20388v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20388v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Compact Vision Models for Iris Presentation Attack Detection under Presentation Attack Instrument Shift and Environmental Degradation</title>
      <link>https://arxiv.org/abs/2609.20386v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20386v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:39:18 GMT</pubDate>
      <dc:creator>Athanasios Angelakis, Marta Gomez-Barrero</dc:creator>
      <category>模型架构</category>
      <description>Iris presentation attack detection (PAD) is security-critical when a subsystem that appears reliable during development encounters presentation attack instruments (PAIs) or acquisition conditions abse...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Compact Vision Models for Iris Presentation Attack Detection under Presentation Attack Instrument Shift and Environmental Degradation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Athanasios Angelakis, Marta Gomez-Barrero</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Iris presentation attack detection (PAD) is security-critical when a subsystem that appears reliable during development encounters presentation attack instruments (PAIs) or acquisition conditions absent from validation data. We benchmark three compact scratch-trained computer-vision models, each with at most approximately 0.26 million trainable parameters, on the Notre Dame subset of LivDet-Iris 2017 under PAI-driven domain shift and environmental degradation. All models are trained without external pretraining or data augmentation and evaluated over five seeds. A validation-selected threshold is transferred unchanged to the known-attack, unknown-attack, corrupted, and pooled test partitions. From known to unknown attack presentations, Attack Presentation Classification Error Rate (APCER) increases by 17.11-30.47 percentage points and Detection Equal Error Rate (D-EER) increases by 7.38-12.73 percentage points. At the validation-selected threshold, ZACH-ViT obtains the lowest unknown-attack APCER (47.69 +/- 4.84%) and D-EER (38.87 +/- 0.93%), while Compact-TransMIL obtains the lowest Bona Fide Presentation Classification Error Rate (BPCER). ZACH-ViT also gives the lowest unknown-attack BPCER at an APCER limit of 10% (81.29 +/- 1.95%). The high absolute errors show that the comparative advantage of the best compact model does not constitute deployment readiness under unknown PAIs.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20386v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20386v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving</title>
      <link>https://arxiv.org/abs/2609.20377v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20377v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:31:59 GMT</pubDate>
      <dc:creator>Shuai Liu, Hechangle Gong, Hao Jiang et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Autonomous driving involves coupled decision-making and scene evolution under multi-mode uncertainty. To capture this coupling and uncertainty, we introduce MM-Future, a world-action model that genera...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shuai Liu, Hechangle Gong, Hao Jiang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Autonomous driving involves coupled decision-making and scene evolution under multi-mode uncertainty. To capture this coupling and uncertainty, we introduce MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair. Each hypothesis is initialized from a structured action prior and an independent future scene source, which are then co-evolved through a modality-aware diffusion Transformer. To support efficient multi-mode rollout, MM-Future compresses multi-view video into planning-oriented representations, dubbed MM-Tokens. Finally, a future-conditioned proposal scorer ranks trajectory candidates by shared history context and their paired predicted future. On NAVSIM navtest, MM-Future achieves 94.0 PDMS and 91.5 EPDMS, while attaining a 32.3 HD-Score in zero-shot closed-loop evaluation on HUGSIM. Ablations show consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20377v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20377v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] [图像生成] Generating Heterogeneous 3D Geological Microstructures from 2D Images via a Stable Diffusion-Adversarial Model</title>
      <link>https://arxiv.org/abs/2609.20358v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20358v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:21:48 GMT</pubDate>
      <dc:creator>Ali Aouf, Eric Laloy, Bart Rogiers et al.</dc:creator>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Characterizing the physical properties of clay and cementitious materials matters across many fields, from materials science to geological waste disposal. Property simulation typically calls for 3D im...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Generating Heterogeneous 3D Geological Microstructures from 2D Images via a Stable Diffusion-Adversarial Model</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ali Aouf, Eric Laloy, Bart Rogiers et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> stable diffusion, gan, denoising diffusion, diffusion, diffusion model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Characterizing the physical properties of clay and cementitious materials matters across many fields, from materials science to geological waste disposal. Property simulation typically calls for 3D imaging, which is expensive, not always accessible, and technically limited for certain materials. Recent progress in deep generative models offers a way around this, reconstructing 3D volumes from the more easily acquired 2D images.   Among GAN-based methods for 3D microstructure generation, SliceGAN has shown strong results for homogeneous isotropic and anisotropic systems. It struggles, however, to capture the finer detail of more complex heterogeneous microstructures, which motivates alternative generative frameworks.   We introduce a hybrid approach that draws on the stability and generation quality of denoising diffusion models. Since no 3D ground truth is available, we replace the standard denoising loss with an adversarial loss, which yields a stable training process in our experiments. We show that the resulting model generates microstructures of varying complexity with minimal slice artefacts and close agreement with ground-truth phase fractions and structural descriptors.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20358v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20358v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts</title>
      <link>https://arxiv.org/abs/2609.20353v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20353v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:17:56 GMT</pubDate>
      <dc:creator>Rui Ai, David Simchi-Levi, Han Zhong</dc:creator>
      <category>模型架构</category>
      <description>We study repeated contract design when a principal observes outcomes but not the actions that generate them. The principal may use any bounded outcome-contingent payment vector, and the agent&apos;s best r...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Rui Ai, David Simchi-Levi, Han Zhong</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We study repeated contract design when a principal observes outcomes but not the actions that generate them. The principal may use any bounded outcome-contingent payment vector, and the agent&apos;s best response can make expected profit discontinuous in those payments. For every fixed number $m\ge2$ of outcomes, the minimax regret over $T$ rounds is of order $T^{m/(m+1)}$, up to logarithmic factors. The upper bound allows arbitrary action spaces and agent heterogeneity, without smoothness or monotone-surplus assumptions. Its key is an effective-dimension reduction that the benchmark can be normalized even when fixed tie-breaking is not shift invariant, after which revealed preference yields a monotone response map in payment-difference coordinates. A learning policy built on a Lipschitz parametrization of this map attains the rate using only observed outcome categories. The lower-bound construction accounts for how incentive losses accumulate across outcome dimensions. It shows that each additional contractible outcome creates a precise and unavoidable increase in the worst-case cost of learning.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20353v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20353v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A Qualitative Model for Reasoning about Path and Support</title>
      <link>https://arxiv.org/abs/2609.20349v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20349v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:13:44 GMT</pubDate>
      <dc:creator>Abhishek Jaiswal, Zoe Falomir</dc:creator>
      <category>模型架构</category>
      <description>Spatial reasoning abilities correlate strongly with performance in STEM fields. Games offer a compelling medium for training these critical skills in developing children who have a natural proclivity ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Qualitative Model for Reasoning about Path and Support</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Abhishek Jaiswal, Zoe Falomir</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Spatial reasoning abilities correlate strongly with performance in STEM fields. Games offer a compelling medium for training these critical skills in developing children who have a natural proclivity for play. However, to facilitate human-like tutoring and player guidance, these games require an AI agent capable of making commonsense inferences from spatial events. Qualitative reasoning (QR) models appear to be a suitable framework for these application domains. As these models reason in symbolic representations, they can seamlessly translate game states into interpretable feedback for human-like player guidance. This paper introduces a hybrid qualitative model designed for Camelot Jr., a block-puzzle game that requires constructing multi-level bridges to connect two avatars stationed on separate towers. The game poses a challenge for the player, who must make platforms stable, plan their path, and ensure they use all the provided blocks. To handle the precise physics required by the domain, we integrate a mathematical center-of-mass stability logic to guide our qualitative solver. Our work facilitates spatial skill training in Camelot Jr. and contributes to the development of human-centric, explainable game-playing agents.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20349v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20349v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] STR-Agent: An LLM-Driven Agent for QoS-Aware Routing in LEO Satellite Networks</title>
      <link>https://arxiv.org/abs/2609.20347v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20347v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:12:22 GMT</pubDate>
      <dc:creator>Bowen Lu, Mugen Peng, Yaohua Sun et al.</dc:creator>
      <category>模型架构</category>
      <description>LEO satellite networks feature dynamic topologies, time-varying links, and diverse service requirements, which make conventional routing schemes difficult to support fine-grained quality-of-service (Q...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">STR-Agent: An LLM-Driven Agent for QoS-Aware Routing in LEO Satellite Networks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Bowen Lu, Mugen Peng, Yaohua Sun et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">LEO satellite networks feature dynamic topologies, time-varying links, and diverse service requirements, which make conventional routing schemes difficult to support fine-grained quality-of-service (QoS) provisioning. Existing studies mainly optimize routing over network states with predefined objectives, but rarely address the practical challenge of translating unstructured natural-language service requests into adaptive routing decisions. To bridge this gap, we propose STR-Agent, an LLM-driven framework for QoS-aware routing in LEO satellite networks. The key innovation of STR-Agent lies in unifying intent perception, tool-based execution, experience accumulation, and reflection-based policy adaptation within a single agent architecture. Specifically, the Perception Module converts natural-language requests into structured routing semantics, while the Reflection Module dynamically adjusts the service-to-routing-policy mapping according to real-time congestion conditions and historical routing outcomes, rather than relying on a fixed routing objective. In addition, we develop a specialized perception model, and construct a domain-specific supervised fine-tuning dataset for LEO service understanding. Simulation results in a Walker-Delta constellation show that STR-Agent significantly outperforms conventional baselines: it reduces end-to-end delay by up to 60% compared with DQ-Dijkstra, improves average intent-understanding accuracy from 45.4% to 92.45% after supervised fine-tuning, and the Reflection Module further reduces the delay by 120 ms at 600 Mbps. These results demonstrate the potential of LLM-driven agent architectures to enable service-aware and adaptive QoS routing in future LEO satellite networks.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20347v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20347v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Fast Cross-Strength Multi-Contrast Brain MRI Translation using Latent Bridge Matching</title>
      <link>https://arxiv.org/abs/2609.20341v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20341v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:10:38 GMT</pubDate>
      <dc:creator>Siddharth Srivastava, Till Bretschneider</dc:creator>
      <category>模型架构</category>
      <description>Magnetic Resonance Imaging (MRI) acquired at different field strengths exhibits pronounced variation in noise, resolution, homogeneity, and contrast, which limits comparability across acquisition sett...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Fast Cross-Strength Multi-Contrast Brain MRI Translation using Latent Bridge Matching</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Siddharth Srivastava, Till Bretschneider</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Magnetic Resonance Imaging (MRI) acquired at different field strengths exhibits pronounced variation in noise, resolution, homogeneity, and contrast, which limits comparability across acquisition settings and complicates downstream analysis. We address this with a unified conditional model for controllable field-to-field synthesis, built on the framework of conditional latent bridge matching. Our single model achieves highly competitive results across the validation phase for all three tasks of the MRIxFields2026 challenge without task-specific architectures or training. We achieve fast generation with only a single inference step, producing all modality and field-strength combinations for $30$ axial slices in under $90$ seconds, as well as cross-modality-strength translation for a full volume in under $70$ seconds, on a single NVIDIA A5000 GPU. We further provide extensive ablations regarding different components of our solution. Code: https://gitlab.com/siddharthsrivastava/mrixfields-2026</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20341v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20341v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Structured Four-Stage Legal Translation: From Natural-Language Traffic Rules to PROLOG</title>
      <link>https://arxiv.org/abs/2609.20334v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20334v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:05:49 GMT</pubDate>
      <dc:creator>May Myo Zin, Wachara Fungwacharakorn, Ken Satoh et al.</dc:creator>
      <category>模型架构</category>
      <description>Traffic regulations are written for human interpretation and therefore rely on shared background knowledge and flexible phrasing, which inherently introduce ambiguity, context dependence, and semantic...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Structured Four-Stage Legal Translation: From Natural-Language Traffic Rules to PROLOG</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> May Myo Zin, Wachara Fungwacharakorn, Ken Satoh et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Traffic regulations are written for human interpretation and therefore rely on shared background knowledge and flexible phrasing, which inherently introduce ambiguity, context dependence, and semantic underspecification. These linguistic characteristics conflict with the precision required by computational reasoning engines such as Prolog, which demand explicit logical structure. This study evaluates two baseline translation approaches, Natural Language to Prolog ($NL\rightarrow Prolog$) and Logical English to Prolog ($LE\rightarrow Prolog$), and introduces a new reasoning-guided translation framework called Structured Four-Stage Legal Translation ($S4L\rightarrow Prolog$). The proposed S4L framework performs semantic role extraction, scene completion, logical mapping, and Prolog rule generation within a single guided prompt, enabling direct translation of raw traffic rules into executable logic without human intervention. A benchmark consisting of twenty real-world traffic rules was used to evaluate each approach in terms of syntactic validity, semantic correctness, and logical completeness. $S4L\rightarrow Prolog$ achieves the highest accuracy, correctly formalizing 75 percent of the rules, while $NL\rightarrow Prolog$ reaches 60 percent and $LE\rightarrow Prolog$ reaches 55 percent. Qualitative analysis further shows that S4L captures implicit causal relations, deontic modality, and exception structure more reliably than the baselines. These results demonstrate that structured reasoning prompts can substantially improve the reliability of natural-language-to-logic translation for legal and safety-critical applications.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20334v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20334v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map</title>
      <link>https://arxiv.org/abs/2609.20333v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20333v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:04:38 GMT</pubDate>
      <dc:creator>Patricia Medina, Hy P. G. Lam</dc:creator>
      <category>模型架构</category>
      <description>We study reconstruction in autoencoders that apply the same forward map before and after setting the observed coordinates to zero. For equal odd input and hidden dimensions $d\geq 3$, among orientatio...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Patricia Medina, Hy P. G. Lam</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We study reconstruction in autoencoders that apply the same forward map before and after setting the observed coordinates to zero. For equal odd input and hidden dimensions $d\geq 3$, among orientation-preserving diffeomorphisms whose Jacobian singular values lie in $[m,M]$, we show that the least uniform reconstruction-derivative error is $\max\{1-M(M-m)/2,0\}$, with affine maps attaining this sharp bound at every prescribed depth. A translated radial rotation can nevertheless reconstruct any prescribed ball exactly with singular values arbitrarily close to one, motivating additional conditions for a finite-data bound. We test this prediction on a 798,452-point terrestrial LiDAR forest scan. At input scale $0.05$, the mean theoretical bound is $0.155$, about $84\%$ of the mean normalized training error $0.185$ across four spatial regions, two depths, and three seeds. At this scale, adding one hidden coordinate reduces the mean reconstruction error below $6\times10^{-6}$.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20333v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20333v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Near-Optimal Pure Single-Loop Extragradient Method for Strongly Convex--Strongly Concave Minimax Optimization</title>
      <link>https://arxiv.org/abs/2609.20327v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20327v1</guid>
      <pubDate>Thu, 17 Sep 2026 13:02:58 GMT</pubDate>
      <dc:creator>Minhao Zhang, Zi Xu</dc:creator>
      <category>模型架构</category>
      <description>We study smooth strongly convex--strongly concave minimax optimization with general nonlinear coupling in the deterministic unconstrained setting. We propose a pure single-loop damped extragradient me...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Near-Optimal Pure Single-Loop Extragradient Method for Strongly Convex--Strongly Concave Minimax Optimization</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Minhao Zhang, Zi Xu</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We study smooth strongly convex--strongly concave minimax optimization with general nonlinear coupling in the deterministic unconstrained setting. We propose a pure single-loop damped extragradient method with fixed parameters and two new full-gradient evaluations per iteration after one initialization query. The method uses an auxiliary feedback recursion and requires no inner solves, accuracy schedules, or staged restarts. We establish last-iterate linear convergence and show that reducing the squared Euclidean distance to the saddle point to an $\varepsilon$ fraction of its initial value requires $O(\sqrt{κ_xκ_y}\log(2κ_xκ_y/\varepsilon))$ full-gradient queries, where $κ_x=L/μ_x$ and $κ_y=L/μ_y$. This bound attains the optimal condition-number order up to logarithmic factors through fixed explicit updates. Numerical experiments demonstrate the effectiveness of the method.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20327v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20327v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images</title>
      <link>https://arxiv.org/abs/2609.20325v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20325v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:59:11 GMT</pubDate>
      <dc:creator>Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain et al.</dc:creator>
      <category>模型架构</category>
      <description>Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Mult...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual grounding capabilities. In this work, we introduce AgriScope, a unified pixel-grounded multimodal framework for agricultural image understanding. AgriScope jointly supports image-level, region-level, and pixel-level understanding within a unified framework, enabling tasks such as grounded caption generation, referring expression segmentation, and multi-turn multimodal interaction for agricultural imagery. AgriScope integrates biologically specialized semantic representations with dense spatial grounding through biological-semantic encoding, dense spatial representations, and pixel decoding. To support large-scale grounded learning, we introduce AgriGround, a large-scale pixel-grounded agricultural multimodal instruction-tuning dataset containing over 500K images and 11M instruction-following samples spanning plant disease analysis, crop and weed identification, insect pest recognition, and fine-grained botanical understanding. AgriGround is constructed through a multi-stage automatic annotation pipeline that integrates multimodal caption generation, phrase-level grounding, segmentation mask generation, and task-oriented instruction synthesis to produce densely grounded supervision. Extensive experiments across multiple agricultural vision-language tasks demonstrate the effectiveness of AgriScope in pixel-grounded multimodal understanding, establishing a strong benchmark for agricultural vision-language learning and visual grounding. The dataset and code will be made publicly available at (https://github.com/boudiafA/AgriScope)</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20325v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20325v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] NeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction</title>
      <link>https://arxiv.org/abs/2609.20323v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20323v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:57:24 GMT</pubDate>
      <dc:creator>Qingde Li, Qingqi Hong, Zihan Li et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Three-dimensional reconstruction from unorganized point clouds remains a challenging problem in computer vision, geometric modeling, and computer-aided design. While neural implicit methods achieve im...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">NeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Qingde Li, Qingqi Hong, Zihan Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Three-dimensional reconstruction from unorganized point clouds remains a challenging problem in computer vision, geometric modeling, and computer-aided design. While neural implicit methods achieve impressive reconstruction accuracy, geometry is typically encoded in latent representations that limit interpretability and reuse within engineering workflows.   We present NeuSOGA3D (Neuro-Symbolic Geometric Abstraction in 3D), a hybrid framework that combines learned perceptual priors inherited from NeuSOGA with explicit symbolic geometric reasoning. The method projects point clouds onto principal orthographic planes, constructs symbolic implicit spline representations from the resulting observations, and fuses them through shape-preserving constructive solid geometry operations to generate a coarse visual hull. Additional geometric detail is recovered through cross-sectional decomposition and volumetric reconstruction using Partial Shape-Preserving Splines.   Unlike conventional neural implicit approaches, NeuSOGA3D progressively transforms observations into explicit symbolic entities, including control polygons, implicit spline fields, cross-sections, and volumetric lofts. Experiments on all forty categories of the ModelNet40 benchmark demonstrate the ability of the framework to recover structurally meaningful and CAD-compatible geometric representations from diverse point-cloud observations. The results highlight the potential of combining learned perception with symbolic geometric reasoning for explainable geometric intelligence.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20323v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20323v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Needles in a Raystack: Ultra-Sparse LiDAR Occupancy Detection for Bat Tracks</title>
      <link>https://arxiv.org/abs/2609.20160v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20160v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:49:42 GMT</pubDate>
      <dc:creator>Nico Klar, Pankaj Rana, Nizam Gifary et al.</dc:creator>
      <category>模型架构</category>
      <description>Monitoring flying animals is important for understanding and protecting biodiversity, but nocturnal species such as bats are difficult to observe in the field. Using LiDAR, bat movements at night resu...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Needles in a Raystack: Ultra-Sparse LiDAR Occupancy Detection for Bat Tracks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Nico Klar, Pankaj Rana, Nizam Gifary et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, u-net</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Monitoring flying animals is important for understanding and protecting biodiversity, but nocturnal species such as bats are difficult to observe in the field. Using LiDAR, bat movements at night result in ultra-sparse 3D spatio-temporal data in which standard reconstruction losses tend to predict only background and miss real flight paths. We study this problem as voxel-wise occupancy detection in sensor-centric LiDAR raystacks. A lightweight 3D U-Net is proposed that preserves temporal resolution, uses skip connections for spatial detail, and combines weighted binary cross-entropy with Dice loss to handle the strong class imbalance.   In real LiDAR recordings of bats over open fields, cross-checked with acoustic monitoring, a reconstruction-based 3D convolutional autoencoder baseline fails to recover foreground trajectories. In contrast, the proposed U-Net recovers sparse foreground occupancy in diagnostic experiments and produces coherent occupancy patterns along bat flight trajectories, providing a practical basis for validation-scale experiments, later clustering of flight tracks, and future integration of bat activity information into biodiversity-aware turbine curtailment strategies.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20160v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20160v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents</title>
      <link>https://arxiv.org/abs/2609.20152v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20152v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:44:03 GMT</pubDate>
      <dc:creator>Pritish Mishra, Ishaan Kumar, Akshat Mandoli et al.</dc:creator>
      <category>模型架构</category>
      <description>Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller&apos;s audio, a language model reads the transcript and decides what to say and which backend tools to call, and...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Pritish Mishra, Ishaan Kumar, Akshat Mandoli et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller&apos;s audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what makes real phone calls hard, such as transcription issues, caller&apos;s voice being split across messages and the requirement that replies follow the language and script specified. We introduce the Multi-Turn Voice Agent Benchmark (MTVA-Bench), which evaluates the language model on the same conditions it faces inside a cascaded system. The caller is played by an LLM following a set of rubrics and tool calls are answered by a mock backend which responds to the arguments the model actually sent. The benchmark contains 49 agents working across 490 reviewed scenarios and supports 7 languages. Scoring is a combination of deterministic checks on tool calls with two LLM judges, one that scores scenario specific rules and one that grades conversation quality without access to the task. Both judges must cite specific messages from the transcript. Task and conversation scores are weighted equally, since a call can complete its task and still go badly for the caller. In a seven-model study, six of the models select the correct tool within 6.4 points of one another, but their overall scores span 24.4 points. Most of the gap comes from argument values, action ordering, rule compliance, and what the model says around its tool calls.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20152v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20152v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Ischemic Stroke Segmentation and Net Water Uptake Quantification on Multicenter Non-Contrast CT Using Supervised Target-Domain Adaptation</title>
      <link>https://arxiv.org/abs/2609.20151v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20151v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:44:00 GMT</pubDate>
      <dc:creator>Linus Britt, Maximilian Nielsen, Susan Klapproth et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Objectives: Quantitative assessment of infarct hypodensity on non-contrast computed tomography (NCCT), including net water uptake (NWU), requires manual or semi-manual lesion delineation, often guided...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Ischemic Stroke Segmentation and Net Water Uptake Quantification on Multicenter Non-Contrast CT Using Supervised Target-Domain Adaptation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Linus Britt, Maximilian Nielsen, Susan Klapproth et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, u-net</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Objectives: Quantitative assessment of infarct hypodensity on non-contrast computed tomography (NCCT), including net water uptake (NWU), requires manual or semi-manual lesion delineation, often guided by CT perfusion or diffusion-weighted MRI, limiting clinical applicability. Automated segmentation on NCCT could enable efficient biomarker extraction such as NWU but remains challenging across heterogeneous multicenter data. This study aimed to develop and externally test a domain-aware deep learning framework for ischemic stroke segmentation on NCCT and assess its suitability for NWU quantification.   Materials &amp; Methods: In this retrospective multicenter study of 801 patients from four datasets, an nnU-Net-based model was trained on NCCT scans from the University Medical Center Hamburg-Eppendorf and the Acute Ischemic Stroke Dataset. To adapt to new domains, the model was fine-tuned on target-domain subsets from Boston (n=11) and ISLES (n=75), with evaluation on held-out cases not used for fine-tuning. Automated segmentations and NWU values were compared with expert references.   Results: For lesions $\geq$ 30 mL, median Dice was 0.68 (Boston) and 0.56 (ISLES). Including smaller lesions, which predominated in ISLES, median Dice was 0.54 (interquartile range [IQR] 0.30-0.70) for acute lesion segmentation (Boston dataset) and 0.20 (IQR 0.03-0.41) for NCCT lesion segmentations when compared to post-treatment infarct (primary target of the ISLES challenge). Automated NWU mean absolute error was 1.37 percentage points (SD 1.61, Boston).   Conclusion: Target-domain adaptation supported NCCT-only infarct segmentation across heterogeneous external cohorts, although performance varied across domains. The approach enabled low-error NWU quantification from baseline NCCT without advanced imaging, supporting further prospective clinical evaluation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20151v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20151v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Task-Oriented Semantic Feature Transmission for Multi-Task Satellite Remote Sensing over Low-SNR Channels</title>
      <link>https://arxiv.org/abs/2609.20150v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20150v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:41:44 GMT</pubDate>
      <dc:creator>Shuoyuan Sun, Hongyu Wang, Mugen Peng et al.</dc:creator>
      <category>模型架构</category>
      <description>Conventional satellite remote sensing transmission follows a reconstruct-then-infer paradigm that optimizes pixel-level fidelity, creating an objective mismatch with downstream tasks such as classific...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Task-Oriented Semantic Feature Transmission for Multi-Task Satellite Remote Sensing over Low-SNR Channels</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shuoyuan Sun, Hongyu Wang, Mugen Peng et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Conventional satellite remote sensing transmission follows a reconstruct-then-infer paradigm that optimizes pixel-level fidelity, creating an objective mismatch with downstream tasks such as classification and detection, especially at low SNR. This paper investigates a task-oriented framework that bypasses image reconstruction and directly transmits semantic features extracted by a multitask-pretrained backbone. A lightweight channel adaptation module (CAM) compresses feature dimensionality for bandwidth reduction, and a feature restorer recovers task-relevant structure after channel corruption. With the backbone frozen, the CAM and task-specific downstream heads are jointly optimized with task and feature-level supervision under random-SNR training. Under the adopted AWGN setting, experiments on scene classification and object detection show consistent gains over reconstruction-oriented JSCC baselines across different SNR conditions, with the largest improvements in the low-SNR regime.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20150v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20150v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Bridging Modalities on the Cortex: Surface-based MRI to PET Translation with a Diffusion Bridge</title>
      <link>https://arxiv.org/abs/2609.20147v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20147v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:38:33 GMT</pubDate>
      <dc:creator>Yitong Li, Alexandra Samoylova, Fabian Bongratz et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Cortical hypometabolism measured by Fluorodeoxyglucose Positron Emission Tomography (FDG-PET) is a highly sensitive biomarker for dementia diagnosis. However, high costs, radiation exposure, and limit...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Bridging Modalities on the Cortex: Surface-based MRI to PET Translation with a Diffusion Bridge</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yitong Li, Alexandra Samoylova, Fabian Bongratz et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Cortical hypometabolism measured by Fluorodeoxyglucose Positron Emission Tomography (FDG-PET) is a highly sensitive biomarker for dementia diagnosis. However, high costs, radiation exposure, and limited accessibility constrain its clinical utility. While cross-modal synthesis from Magnetic Resonance Imaging (MRI) offers a promising alternative, existing volumetric generation methods do not explicitly account for the highly folded cortical geometry, where disease-related patterns predominantly reside. To address this, we introduce a novel surface-based diffusion bridge framework DB-SUiT for MRI-to-PET translation that operates natively on the cortical manifold. A conditional Spherical U-shaped vision Transformer (SUiT) is specifically designed to model the intricate cross-modal relationships while preserving surface topology. It combines spherical convolutional encoders for multi-scale surface feature extraction with bottleneck Transformers to capture long-range spatial dependencies, while incorporating demographic and subcortical conditions to refine the synthesis. Evaluated on two datasets, including subjects with different dementia types, DB-SUiT demonstrates high-fidelity synthesis that substantially outperforms other baselines. In automated dementia classification, synthesized PET surfaces improve performance over MRI by 14.2% and PET volumes by 11.3%, approaching the performance of real PET surfaces. In a blinded reader study, synthetic PET achieved 85.5% diagnostic accuracy, compared with 75.8% for MRI and 95.2% for real PET. This further demonstrates cross-cohort and cross-pathology generalization, as the model was evaluated without retraining on an external cohort that included a dementia subtype not represented during training. Our code is available at https://github.com/ai-med/DB-SUiT.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20147v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20147v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models</title>
      <link>https://arxiv.org/abs/2609.20139v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20139v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:31:07 GMT</pubDate>
      <dc:creator>Farooq Ahmad Wani, Maria Sofia Bucarelli, Mujtaba Hussain Mirza et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Vision-language models (VLMs) are fragile under image corruption. We find that the wording of the question affects VLMs in two opposite ways. Verbose questions make VLMs substantially more robust---e....</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Farooq Ahmad Wani, Maria Sofia Bucarelli, Mujtaba Hussain Mirza et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Vision-language models (VLMs) are fragile under image corruption. We find that the wording of the question affects VLMs in two opposite ways. Verbose questions make VLMs substantially more robust---e.g., rephrasing &quot;Is there a cat?&quot; into &quot;Please look carefully and answer: is there a cat?&quot;. Conversely, VLMs become more fragile under corruption when the question is semantically complex or finer-grained, e.g., &quot;what colour is the cup left of the chair?&quot; instead of &quot;is there a cup?&quot;. Both effects stem from question-conditioned cross-modal attention, which induces a spectral filter over image patches: verbose questions broaden its frequency support, while fine-grained questions concentrate it onto fewer visual scales. The model&apos;s answer drifts most when this filter and the corruption sit on the same spatial frequencies. We test the filter view on Qwen3-VL and LLaVA-OneVision across GQA and CLEVR; verbose paraphrasing reduces drift variance by 70--81% on the 8B models. The practical recipe---pad the prompt---further yields measurable gains in accuracy, even under image corruption.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20139v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20139v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Fast-varying Natural Frequencies and Damping Ratio Identification for Linear Time-Varying System</title>
      <link>https://arxiv.org/abs/2609.20138v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20138v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:29:44 GMT</pubDate>
      <dc:creator>Melisa Bozaci, Alice Cicirello</dc:creator>
      <category>模型架构</category>
      <description>This work proposes a physics-enhanced machine learning approach for the system identification of Linear Time-Varying (LTV) systems under time-varying operating conditions in terms of fast-varying natu...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Fast-varying Natural Frequencies and Damping Ratio Identification for Linear Time-Varying System</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Melisa Bozaci, Alice Cicirello</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">This work proposes a physics-enhanced machine learning approach for the system identification of Linear Time-Varying (LTV) systems under time-varying operating conditions in terms of fast-varying natural frequencies and damping ratios by combining a long short-term memory network with an Extended Kalman Filter (EKF). The proposed approach uses vibration data (displacement and velocity measurements), domain knowledge of modal damping ratios, and a physics-based model that can yield an approximate natural frequencies time-dependency model. The approach is validated using synthetic data generated from a finite element model of a 2-blade offshore wind turbine under realistic environmental and operating conditions. This system displays fast time-varying frequencies due to operating conditions, whose identification is particularly challenging because of the wind and wave loading. The robustness of the proposed approach is assessed under assumed incorrect system information (e.g. damping ratio). The proposed approach is evaluated across different environmental and operating conditions to show its applicability to different operating regimes. The results show the approach can accurately identify the selected fast-varying natural frequency, 1st Fore-Aft (FA-1) mode, with a maximum root mean square error of 0.0012 Hz. The results demonstrate that the model trained on EKF estimates depends on accurate damping values, whereas the model trained on physics-based data exhibits robustness to incorrect damping assumptions. The approach is extended to damping ratio identification for the selected mode by estimating the root mean square error between models trained on EKF estimates and physics-based data. The results show that the approach can yield a good approximation of the FA-1 mode damping ratio using grid search, offering an improvement over covariance-driven stochastic subspace identification.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20138v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20138v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis</title>
      <link>https://arxiv.org/abs/2609.20124v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20124v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:20:48 GMT</pubDate>
      <dc:creator>Zifan Guan, Longyu Lu, Junan Zhang et al.</dc:creator>
      <category>模型架构</category>
      <description>Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. Wh...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zifan Guan, Longyu Lu, Junan Zhang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspects, they are too costly for massive inference and reinforcement learning feedback. To address this, we first introduce Live-ProsodyJudge (LPJ), a cost-effective pairwise evaluator distilled from Gemini into Qwen3-Omni. However, we identify a critical flaw in standard multi-dimensional evaluation: verdict coupling. The judge tends to lazily align all individual dimension scores with its overall preference, collapsing a rich multi-dimensional rubric into a single preference bit. To resolve this, we further propose Decoupled-Live-ProsodyJudge (D-LPJ). D-LPJ eliminates the overall verdict target to prevent blind following, masks uncertain pair-dimensions during Supervised Fine-Tuning(SFT), and introduces a novel span-local GRPO strategy that applies normalized advantages strictly to their corresponding rationale spans. Evaluated on highly curated human-annotated test sets, 10 sample balanced-order LPJ achieves higher point accuracy than a single Gemini call, while D-LPJ successfully produces independent,decoupled dimension judgments. Furthermore, in a Best-of-8 TTS candidate selection tournament, the LPJ-selected utterance falls within the human top-3 in 85.29% of high-confidence cases, demonstrating its efficacy for fine-grained TTS preference optimization.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20124v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20124v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents</title>
      <link>https://arxiv.org/abs/2609.20110v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20110v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:09:21 GMT</pubDate>
      <dc:creator>Yichao Jin, Yushuo Wang, Yuxuan Han et al.</dc:creator>
      <category>多模态生成</category>
      <description>Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yichao Jin, Yushuo Wang, Yuxuan Han et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting the key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with a final conformal risk control, the score can be used for reliable STP of financial documents. The method is validated on three public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 for VLM verbalized signals to 0.90-0.99 with contributions from all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence signals could clear only 0.1%-7.0% of fields under risk control at a target error of &lt;10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error of the accepted tier at or below the target.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20110v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20110v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[评估与优化] [模型架构] AnyviewMeter: Adapting Robotic Reward Models with Camera Geometry and Multi-View Attention</title>
      <link>https://arxiv.org/abs/2609.20106v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20106v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:07:22 GMT</pubDate>
      <dc:creator>Yuang Tu, Runjia Tan, Yujie Yan et al.</dc:creator>
      <category>评估与优化</category>
      <category>模型架构</category>
      <description>Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AnyviewMeter: Adapting Robotic Reward Models with Camera Geometry and Multi-View Attention</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yuang Tu, Runjia Tan, Yujie Yan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 评估与优化, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> reward model, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting a pretrained reward model to a local task therefore requires accounting for how that task is observed. We introduce AnyviewMeter, a geometry-conditioned adaptation framework for robotic reward models that represent task progress as a scalar reward signal. It combines low-rank fine-tuning with token-aligned Plucker rays and synchronous block attention: ray conditioning incorporates camera geometry into visual features and attention queries and keys, while block attention fuses synchronized views inside the pretrained decoder. The framework supports both single-view reward prediction and joint multi-view evaluation through parameter-efficient adaptation of a pretrained Robometer model. On PickCube, single-view adaptation improves progress prediction in every camera group and reduces mean absolute error under a changed field of view by approximately 21% relative to RGB fine-tuning. Across simulated manipulation tasks, joint multi-view prediction reduces progress error by 41-69% compared with averaging single-view RGB predictions and improves temporal ordering in approximately 88% of task-camera groups. On real tasks with fixed and wrist-mounted cameras, mean absolute error decreases by approximately 21% relative to averaged RGB fine-tuning. These results support camera geometry and joint visual evidence as useful components of task-specific robotic reward adaptation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20106v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20106v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 评估与优化, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A Smaller Transformer in Your Transformer</title>
      <link>https://arxiv.org/abs/2609.20100v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20100v1</guid>
      <pubDate>Thu, 17 Sep 2026 12:01:13 GMT</pubDate>
      <dc:creator>Dhananjay Tomar, Marius Aasan, Andreas Kleppe et al.</dc:creator>
      <category>模型架构</category>
      <description>Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this re...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Smaller Transformer in Your Transformer</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Dhananjay Tomar, Marius Aasan, Andreas Kleppe et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20100v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20100v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting</title>
      <link>https://arxiv.org/abs/2609.20086v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20086v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:45:32 GMT</pubDate>
      <dc:creator>Abraham Ezema, Chijioke Eze, Ferdinanda Ponci et al.</dc:creator>
      <category>模型架构</category>
      <description>Long-term multivariate time series plays a significant role in many application areas such as power systems, trading, etc. However, their accurate prediction is quite difficult for conventional foreca...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Abraham Ezema, Chijioke Eze, Ferdinanda Ponci et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Long-term multivariate time series plays a significant role in many application areas such as power systems, trading, etc. However, their accurate prediction is quite difficult for conventional forecasting methods as they often exhibit high dimensionality and complex relationships. Recent works show that transformer-based approaches are quite effective for long-term forecasting thanks to their attention mechanism. However, in the presence of complex high-dimensional inputs, they show evidence of oversmoothing, limited capacity, and opacity. To this end, this paper introduces SETTer, a transformer-based model that addresses these challenges by incorporating novel techniques for decoupled self-attention and hybrid masking. The proposed techniques enable SETTer to effectively capture the dominant short- and long-term patterns across the temporal and channel dimensions. In addition, we enrich the model layers with simple explainable structures that indicate the discriminative pattern of SETTer. We show that with a single-layer transformer architecture, SETTer can effectively model long-term dependencies in the presence of varying data complexities. Extensive experiments on real-word benchmark datasets for long-term multivariate time series forecasting demonstrate that SETTer outperforms state-of-the-art models in 88% of the scenarios.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20086v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20086v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards</title>
      <link>https://arxiv.org/abs/2609.20082v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20082v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:40:26 GMT</pubDate>
      <dc:creator>Shihao Liu, Hao Yin, Lijun Liu et al.</dc:creator>
      <category>模型架构</category>
      <description>Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current method...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shihao Liu, Hao Yin, Lijun Liu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current methods still face two problems: fixed-threshold curricula can become misaligned with the policy&apos;s evolving capability boundary, and additive rewards can leak argument-level credit when the predicted tool is wrong. To address these problems, we propose MATCH, a closed-loop framework for model-aware tool learning with curriculum scheduling and hierarchically gated rewards. Model-Aware Curriculum Learning (MACL) maintains reward-derived sample difficulty that co-evolves with the policy, and each epoch selects samples near the current capability boundary together with a top-k pool of harder cases. Hierarchical Tool-call Gated Reward (HTGR) scores tool name, argument key, and argument value as a gated chain, granting credit at each level only when prerequisites hold. The same HTGR rewards drive both GRPO updates and MACL&apos;s difficulty refresh, closing the loop between policy optimization and sample scheduling. On API-Bank and BFCL V3, MATCH reaches 72.19% and 62.87% overall accuracy, outperforming the main supervised and RL-based baselines. Backbone experiments further show consistent improvements across four backbones from two model families.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20082v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20082v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Tailored to you: longitudinal effects of personalising language models</title>
      <link>https://arxiv.org/abs/2609.20077v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20077v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:35:06 GMT</pubDate>
      <dc:creator>Canfer Akbulut, Justine Breuch, Arianna Manzini et al.</dc:creator>
      <category>模型架构</category>
      <description>Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions w...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Tailored to you: longitudinal effects of personalising language models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Canfer Akbulut, Justine Breuch, Arianna Manzini et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people&apos;s perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users&apos; self-perceptions and interpersonal relationships, remain largely unexamined. In this study, we recruited 992 participants to complete daily advice-seeking interactions with language models over the course of five days, comparing outcomes from a non-personalised baseline against two personalisation approaches: memory-based (conditioned on prior conversational history) and survey-based (conditioned on information collected through a pre-study intake survey). We find that several changes in human-AI interaction over time are driven primarily by repeated exposure rather than personalisation itself. However, participants interacting with personalised models experienced differences in advice-seeking and information-sharing attitudes and behaviours: participants in the memory-based condition engaged in greater self-disclosure and rated the model as less creepy, while participants in the survey-based condition reported higher regret about having shared personal information with the AI. We conclude by highlighting the nuanced effects of different personalisation approaches on interaction outcomes, and discussing the implications of these findings for the responsible design and deployment of personalised AI systems.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20077v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20077v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference</title>
      <link>https://arxiv.org/abs/2609.20068v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20068v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:19:58 GMT</pubDate>
      <dc:creator>Caroline Gans Combe</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key--Value cache of transformer language models...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Caroline Gans Combe</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, lora, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key--Value cache of transformer language models. The singular value spectrum of a rating matrix is shown to be a diminishing marginal utility schedule for latent factors, the eigenvalue spectrum of the projected covariance operator to be the marginal utility schedule of a model&apos;s learned representation, and cache eviction and low-rank cache compression to be instances of constrained utility maximization under a memory budget. The three collapse into a single allocation rule: retain the top dimensions whose eigenvalue exceeds the shadow price of the binding constraint. The framework is applied to the automated extraction of structured information from geo-mining documents, where it motivates a multi-pass inference protocol, a layer-wise TIES model merging procedure, and a selection policy combining extraction quality, localization drift and energy, scalarized with a Conditional Value-at-Risk term on drift. Two empirical contributions are reported. An 11.2-million-parameter hierarchical classifier, trained in about five minutes on a single GPU, reaches 90.0 per cent level-1 accuracy on a held-out test set from a 973-document uranium-exploration corpus, against 92.0 per cent for a proprietary model on a fifty-document human audit of the same corpus, at a latency of 2.62 ms per card against approximately 2,000 ms for the API and at negligible cost. A diagnostic of uniform-density TIES merging exposes a reproducible degenerate mode in which the merged model returns token-identical outputs across five geographically distinct districts while declaring high confidence; re-executing the merge under layer-wise calibrated densities removes that signature on the diagnostic sample. The full-scale extraction benchmark, including LoRA fine-tuning, is reported as projected rather than measured and remains an empirical extension of this work.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20068v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20068v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity</title>
      <link>https://arxiv.org/abs/2609.20067v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20067v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:19:26 GMT</pubDate>
      <dc:creator>Abdullahi Isa, Souley Boukari, Muhammad Aliyu</dc:creator>
      <category>模型架构</category>
      <description>Deep learning models for multi-modal breast cancer diagnosis achieve high predictive accuracy but remain clinically unacceptable without actionable, counterfactual explanations. Attribution-based meth...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Abdullahi Isa, Souley Boukari, Muhammad Aliyu</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Deep learning models for multi-modal breast cancer diagnosis achieve high predictive accuracy but remain clinically unacceptable without actionable, counterfactual explanations. Attribution-based methods (LIME, SHAP) are categorically inapplicable to this purpose, as they generate no alternative instances and thus cannot be evaluated on counterfactual quality metrics. This investigation provides empirical evidence that FCA-Guided Counterfactual (FCA-CF) framework that uses a Formal Concept Analysis (FCA) concept lattice as a hard structural constraint on counterfactual search, operating over a multi-modal TCGA-BRCA dataset. We benchmark against four genuine counterfactual methods: Wachter-style CF, DiCE, FACE, and NICE, evaluated on 60 benign-predicted TCGA-BRCA instances. The FCA-CF framework achieves Validity = 1.0000 (100% of counterfactuals successfully flip the prediction), Sparsity = 2.37 features changed (best among all valid methods), and Proximity = 0.900 (normalised L2-based, matching NICE as joint best). The classifier achieves Accuracy = 0.980, F1 = 0.976, ROC-AUC = 0.9947. Ablation analysis confirms that the FCA lattice constraint is the primary sparsity driver (removing it increases sparsity by +40%, p &lt; 0.001, Cohen&apos;s d = 0.78), while Phase C greedy refinement accounts for the largest individual contribution (+113% sparsity increase when disabled, p &lt; 0.001, d = 5.01). FCA-guided counterfactual generation achieves a clinically important Pareto-dominant outcome; it is simultaneously the sparsest and among the most proximate of all valid methods, with perfect validity. The emergent sparsity property arising from lattice topology rather than numerical penalty terms constitutes a structurally novel contribution to the counterfactual explanation literature.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20067v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20067v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation</title>
      <link>https://arxiv.org/abs/2609.20066v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20066v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:16:59 GMT</pubDate>
      <dc:creator>Zongze Wu, Baofeng Jia, Weiqi Yan et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Event cameras offer high temporal resolution and motion sensitivity for tiny UAV detection, yet distant targets generate sparse and fragmented events that are easily overwhelmed by clutter and ego-mot...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zongze Wu, Baofeng Jia, Weiqi Yan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Event cameras offer high temporal resolution and motion sensitivity for tiny UAV detection, yet distant targets generate sparse and fragmented events that are easily overwhelmed by clutter and ego-motion. Existing methods mainly rely on dense event representations or local sparse spatiotemporal modeling, resulting in redundant computation or fragmented modeling of motion continuity across distant asynchronous events. To address this limitation, we introduce serialized motion evidence accumulation, which treats motion continuity as an ordered evidence propagation process. Specifically, the same event stream is organized into locality-preserving spatiotemporal paths and chronology-preserving temporal paths through the latent complementary serializations. Based on this principle, we propose PointEvent, a lightweight event-wise state-space framework that alternates serialized scans across the complementary orders, progressively consolidating fragmented motion evidence beyond fixed local neighborhoods. A high-resolution event branch preserves fine-grained target responses, while compact context modulation suppresses interference. Experiments demonstrate that PointEvent achieves SOTA with the fewest parameters and fastest measured inference among the compared methods. Code: https://github.com/wzz-z/PointEvent</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20066v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20066v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition</title>
      <link>https://arxiv.org/abs/2609.20064v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20064v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:15:07 GMT</pubDate>
      <dc:creator>Benjamin Kiessling</dc:creator>
      <category>多模态生成</category>
      <description>Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-sc...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Benjamin Kiessling</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-scale pretraining, and hallucination. Historical ATR therefore continues to rely largely on compact CRNN line recognizers, which are visually grounded and trainable on modest data. Lightweight recurrence-free recognizers promise the accuracy of larger models with the practical advantages of CRNNs, yet have not been comprehensively evaluated on historical writing. We adapt PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compare it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- and Arabic-script material. While PP-OCRv6 does not consistently outperform the baseline when trained from scratch, heterogeneous pretraining produces markedly better generalization. Comparisons with the Qwen3.5-based Medusa recognizer further show that fine-tuned PP-OCRv6 can outperform a large VLM tailored towards historical Latin-script HTR.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20064v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20064v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection</title>
      <link>https://arxiv.org/abs/2609.20063v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20063v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:14:58 GMT</pubDate>
      <dc:creator>Xiang Li, Pin-Yu Chen, Wenqi Wei</dc:creator>
      <category>模型架构</category>
      <description>The rapid advancement of speech synthesis and voice conversion technologies has made audio deepfakes increasingly realistic, posing serious security risks in practical applications. While existing det...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xiang Li, Pin-Yu Chen, Wenqi Wei</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The rapid advancement of speech synthesis and voice conversion technologies has made audio deepfakes increasingly realistic, posing serious security risks in practical applications. While existing detection methods achieve strong performance under controlled conditions, they often fail to generalize under real-world perturbations and corruptions. In this paper, we propose ROGUE, a framework that dynamically constructs robust detection workflows by orchestrating multiple detection tools. ROGUE formulates workflow generation as a sequential decision-making problem and introduces a dual-agent paradigm, where a perturbation agent generates audio perturbations and a policy agent learns to select and execute detection tools under perturbed conditions. Through adversarial learning, ROGUE enables perturbation-aware tool selection, adaptive execution strategies, and improved robustness to distribution shifts. Extensive experiments across multiple datasets and real-world corruptions demonstrate that ROGUE consistently outperforms strong baselines in both robustness and generalization. Our results highlight the effectiveness of adversarially optimized workflow generation for building reliable audio deepfake detection systems in real-world deployment settings.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20063v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20063v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Evaluating Explanation Methods by the Predictors They Induce</title>
      <link>https://arxiv.org/abs/2609.20058v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20058v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:10:19 GMT</pubDate>
      <dc:creator>Jacob Selbæk, Hugo L. Hammer</dc:creator>
      <category>模型架构</category>
      <description>Explanations of machine learning models are usually judged by criteria that are hard to compare. We propose a simpler test: if an explanation really describes how a model uses its features, it should ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Evaluating Explanation Methods by the Predictors They Induce</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jacob Selbæk, Hugo L. Hammer</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Explanations of machine learning models are usually judged by criteria that are hard to compare. We propose a simpler test: if an explanation really describes how a model uses its features, it should be possible to rebuild the model&apos;s predictions from it. We turn each explanation into a predictor by reading each feature&apos;s effect and adding them up, and measure how well that predictor reproduces the model on unseen data. Nothing is fitted, so the score reflects the explanation itself. The test applies to any explanation that can be written as a function of the features; we demonstrate it on partial dependence plots (PDP), accumulated local effects (ALE), SHAP and LIME. We prove that summing partial dependence curves gives the best possible additive summary of a model when its features are independent, and that this fails when they are dependent. Across 13 real datasets and 9 synthetic designs and four model families, which method scores best depends entirely on feature dependence: where features are independent SHAP is slightly worse than PDP, exactly as the theory predicts; on dependent real data SHAP leads. Some widely used quality metrics even prefer a damaged explanation to an intact one.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20058v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20058v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] WiCleanData: Guaranteeing the Type Consistency of Wikidata by Taxonomy Refinement and Constraint Enforcement</title>
      <link>https://arxiv.org/abs/2609.20057v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20057v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:09:51 GMT</pubDate>
      <dc:creator>Yiwen Peng, Marc Jeanmougin, Thomas Bonald</dc:creator>
      <category>图像生成</category>
      <description>Because of its collaborative nature, Wikidata suffers from errors, in- consistencies, and excessive complexity, such as redundant classes, ambiguity between instances and classes, wrong taxonomic path...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">WiCleanData: Guaranteeing the Type Consistency of Wikidata by Taxonomy Refinement and Constraint Enforcement</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yiwen Peng, Marc Jeanmougin, Thomas Bonald</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Because of its collaborative nature, Wikidata suffers from errors, in- consistencies, and excessive complexity, such as redundant classes, ambiguity between instances and classes, wrong taxonomic paths, and type constraint violations. The manual curation of these issues is infeasible at scale. To address these challenges, we introduce WiCleanData, a refined version of Wikidata with a consistent tax- onomy and free from type constraint violations. Specifically, we have designed an automated pipeline that first cleans the taxonomy with language model assistance, then simplifies type constraints by hierarchical aggregation, and finally filters facts accordingly. The resulting knowledge graph, free from any type violation, is made publicly available via a Web interface, enabling easy exploration and downstream applications.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20057v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20057v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution</title>
      <link>https://arxiv.org/abs/2609.20056v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20056v1</guid>
      <pubDate>Thu, 17 Sep 2026 11:07:47 GMT</pubDate>
      <dc:creator>Loan Bernat, Matthieu Grard, Ariane Herbulot et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Hierarchical robotic systems executing long-horizon manipulation tasks must make high-level semantic decisions that orchestrate stochastic low-level skills. In this setting, failed rollouts are ambigu...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Loan Bernat, Matthieu Grard, Ariane Herbulot et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Hierarchical robotic systems executing long-horizon manipulation tasks must make high-level semantic decisions that orchestrate stochastic low-level skills. In this setting, failed rollouts are ambiguous: a poor downstream state may reflect an invalid high-level decision, partial observation, or a valid decision whose physical execution failed. Traditional supervised learning lacks data for such recovery states, while reinforcement learning struggles with sparse rewards and non-local credit assignment. We propose MAGMA-GEN, an on-policy data-generation pipeline that converts ambiguous failed rollouts into validated recovery supervision. MAGMA-GEN first uses a privileged coach to hypothesize an early decision-level error and propose localized correction or recovery actions. Because this diagnosis is fallible, candidates are retained only if re-execution from the same state under matched conditions improves downstream progress. This produces supervised examples from the agent&apos;s own failure distribution without per-step human demonstrations. Evaluated on interactive long-horizon manipulation tasks, MAGMA-GEN improves task success and recovery capabilities, against distillation and trajectory-repair baselines under evolving task constraints in both simulation and real-robot execution.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20056v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20056v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [模型架构] [扩散模型] [图像生成] DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models</title>
      <link>https://arxiv.org/abs/2609.20051v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20051v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:58:04 GMT</pubDate>
      <dc:creator>Shihong Li, Juntao Xu, JinCao et al.</dc:creator>
      <category>视频生成</category>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Step distillation reduces the cost of video generation, but reusing a LoRA trained for a longer trajectory can alter its functional effect or degrade target quality. Static parameter compatibility off...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shihong Li, Juntao Xu, JinCao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> video generation, lora, distillation, dit, diffusion</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Step distillation reduces the cost of video generation, but reusing a LoRA trained for a longer trajectory can alter its functional effect or degrade target quality. Static parameter compatibility offers one perspective on this problem; our observations show that similar measured geometry can coexist with different adapter behavior under a shortened denoising schedule. We propose DART, a training-free method that combines low-rank coordinate transport with target-schedule response calibration using forward evaluations and no source training videos. On a four-step Wan2.2 target, DART-F improves the joint quality score from 0.9029 to 0.9227 and changes macro functional retention from -0.4644 to +0.1349. Component analysis shows that calibration accounts for most of the quality improvement, while coordinate transport provides complementary gains when combined with calibration. Adapter-level results reveal positive functional effects for some adapters and strong attenuation with reduced negative functional effects for others. Evaluations on two additional targets show the same aggregate trend. These results motivate evaluating distilled-model LoRA reuse jointly through functional preservation and negative-transfer avoidance, without assuming recovery for every adapter.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20051v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20051v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents</title>
      <link>https://arxiv.org/abs/2609.20050v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20050v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:56:47 GMT</pubDate>
      <dc:creator>Zhexi Feng, Ruiyi Zhang, Yongbo Yang et al.</dc:creator>
      <category>模型架构</category>
      <description>A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhexi Feng, Ruiyi Zhang, Yongbo Yang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with variants of one required fact and leave the decision unsupported. We formulate state-conditioned minimal sufficient evidence recovery: given a captured agent state, recover a compact evidence combination that supplies the support its next decision still lacks. SERBench measures this on 500 held-out states from 45 repositories, recording what the agent has seen and crediting only sets that cover every fact the current decision was annotated to require. MSS-Complement treats acquisition as set construction, not ranking. Three semantic calls propose a jointly sufficient set, search for what it lacks, and return 4-8 intact source units within 6,144 tokens. One configuration, fixed on calibration data, recovers a complete set for 73.0% of those states at five items and 80.6% at eight, against 61.4% and 72.4% for Qwen3 embedding with reranking. A matched control ranking by similarity alone reaches 66.6%, placing the gain in the set-level policy, not the computation. From frozen repository source with no gold-derived pool, the lead is 5.0 points. On AMA-Bench it answers from a 76.2% smaller answer prompt, with accuracy 2.08 points above that benchmark&apos;s own memory agent. Removing one required group from an otherwise complete set costs 12.3 and 11.1 points of repair-localization precision under two executors. Retrieval for agents is better posed as recovering what a decision lacks than re-ranking what an issue resembles.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20050v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20050v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression</title>
      <link>https://arxiv.org/abs/2609.20045v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20045v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:51:37 GMT</pubDate>
      <dc:creator>Guangzhe Zhang</dc:creator>
      <category>模型架构</category>
      <description>A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current ans...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Guangzhe Zhang</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers. A pilot evaluates 24 history pairs across six synthetic mechanisms, 12 memory conditions, two repeats, and two model backends. A deterministic frontier selector obtains strict reveal accuracy of 96/96 on DeepSeek and 82/96 on GLM; a structured writer obtains 62 successes with one unresolved outcome and 56/96. The configured four-outcome joint contrast has finite-sample identification intervals of [0.521, 0.542] and [0.292, 0.313], not confidence intervals. A record-level audit distinguishes retained-state adequacy, response delivery, and answer-schema compliance without changing those original scores. It finds 26 and 25 well-formed but semantically wrong structured reveal memories, while all 14 GLM frontier reveal failures contain correct values in the wrong wrapper. Tombstone removal produces 16/16 exact replay failures in the targeted mechanism. Identifier renaming then exposes a separate flaw: original frontier late-reference adequacy falls from 8/8 to 94/320 transformed instances. We provide and test a label-equivariant repair, but it preserves only 2/8 original late-reference answers: eliminating a naming shortcut does not solve unknown future relevance. These results support a scoped evaluation methodology and reproducible failure analysis, not general superiority of the repaired algorithm. Paid pilot evidence, retrospective diagnostics, and new offline tests are reported separately; no independent held-out or natural-task validation is claimed.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20045v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20045v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [模型架构] Astronex-World 1.0: Real-Time Interactive World Model Foundation</title>
      <link>https://arxiv.org/abs/2609.20034v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20034v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:38:22 GMT</pubDate>
      <dc:creator>Xin Zhou, Cong Miao</dc:creator>
      <category>视频生成</category>
      <category>模型架构</category>
      <description>We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual state...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Astronex-World 1.0: Real-Time Interactive World Model Foundation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xin Zhou, Cong Miao</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> text-to-video, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20034v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20034v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Governance-as-Code: Translating EU AI Act Technical Requirements into Executable Compliance Pipelines for Generative AI Systems</title>
      <link>https://arxiv.org/abs/2609.20016v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20016v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:23:54 GMT</pubDate>
      <dc:creator>Rudrendu Kumar Paul, Sourav Nandy</dc:creator>
      <category>模型架构</category>
      <description>The EU AI Act (Regulation 2024/1689) imposes technical obligations on high-risk AI providers, yet Articles 8-15 were drafted for predictive AI and leave seven technical gaps when applied to generative...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Governance-as-Code: Translating EU AI Act Technical Requirements into Executable Compliance Pipelines for Generative AI Systems</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Rudrendu Kumar Paul, Sourav Nandy</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The EU AI Act (Regulation 2024/1689) imposes technical obligations on high-risk AI providers, yet Articles 8-15 were drafted for predictive AI and leave seven technical gaps when applied to generative systems, spanning non-deterministic data governance, training-data provenance, continuous conformity, human oversight, open-ended robustness, emergent risk, and generative fairness. We deliver Governance-as-Code (GaC), a framework of 43 machine-checkable acceptance criteria across six compliance modules that run in a CI/CD pipeline and emit Article-indexed audit evidence, and we show the actual Rego policy code rather than merely describing it. Our central commitment is that the Act&apos;s open-textured standards (&quot;appropriate levels,&quot; &quot;possible biases&quot;) become declared, auditable numbers: robustness thresholds are derived from the provider&apos;s documented baseline and a state-of-the-art floor, and framing bias is collapsed into eight measurable proxies tested by counterfactual demographic probing. We also correct who owes what, since under Article 25 and Chapter V a downstream deployer relies on the upstream provider&apos;s Article 53 training-data summary and documents only the layers it controls, so GaC verifies that summary rather than demanding per-sample documentation the deployer never had. We validate on two enterprise deployments, a high-risk advisory chatbot and a limited-risk content generator, benchmarking against a manual expert audit rather than documentation artifacts that were never designed to enforce compliance. GaC reproduces all of the manual audit&apos;s findings, including three penalty-triggering violations, while cutting audit labor by roughly 75%.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20016v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20016v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction</title>
      <link>https://arxiv.org/abs/2609.20012v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20012v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:21:52 GMT</pubDate>
      <dc:creator>Enpeng Li, Yunzhou Zhang, Zhiyao Zhang et al.</dc:creator>
      <category>图像生成</category>
      <description>Feed-forward 3D reconstruction provides an efficient paradigm for scene modeling from image sequences. Scaling these models to large monocular scenarios are constrained by excessive GPU memory footpri...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Enpeng Li, Yunzhou Zhang, Zhiyao Zhang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Feed-forward 3D reconstruction provides an efficient paradigm for scene modeling from image sequences. Scaling these models to large monocular scenarios are constrained by excessive GPU memory footprint, degraded local geometry, and long-term trajectory drift. Existing chunk-based optimization strategies provide limited geometric constraints and fail to maintain global consistency over extended trajectories. We present a unified framework for stable and scalable feed-forward 3D reconstruction from long monocular sequences. Our approach builds on coarse-to-fine trajectory alignment augmented by lightweight geometric prior injection. Distilling monocular geometric cues into the feed-forward backbone via LoRA adaptation improves depth accuracy on fine structures while preserving inference efficiency. We introduce a hybrid-weight sparse ray-field optimization that leverages high-frequency geometric features to guide local point-cloud refinement and enforce consistent inter-frame ray constraints. Unlike prior chunk-based methods, this establishes strong cross-frame geometric coupling while maintaining scalability. Finally, an efficient trajectory stitching strategy with joint ray-error optimization explicitly reduces accumulated drift. Extensive experiments show that our approach achieves competitive trajectory accuracy compared with representative SLAM systems, while maintaining globally consistent 3D reconstruction in large-scale scenarios.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20012v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20012v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Dynamic Generalized Gromov-Wasserstein Optimal Transport</title>
      <link>https://arxiv.org/abs/2609.20008v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20008v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:17:33 GMT</pubDate>
      <dc:creator>Junda Ying, Zhiwei Zeng, Peijie Zhou et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Gromov--Wasserstein optimal transport (GW-OT) extends classical optimal transport by introducing structure-aware transport cost. This is particularly relevant for spatial transcriptomics, where dynami...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Dynamic Generalized Gromov-Wasserstein Optimal Transport</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Junda Ying, Zhiwei Zeng, Peijie Zhou et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> flow matching, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Gromov--Wasserstein optimal transport (GW-OT) extends classical optimal transport by introducing structure-aware transport cost. This is particularly relevant for spatial transcriptomics, where dynamical reconstruction should preserve tissue structure in addition to matching expression patterns. While static formulations have been widely used for such structure-aware alignment, a general dynamic formulation for reconstructing continuous trajectories is still missing. We introduce Travelling Pair Dynamical Alignment and Trajectory Estimation (TP-DATE), a theoretical and computational framework to generalize GW-OT dynamically in a simulation-free manner. We formulate a broad class of static and dynamic Quadratic-form OT (QOT) through path actions and prove the static dynamic equivalence. We further develop travelling-pair flow matching, which allows interacting conditional paths and marginalizes their interactions into a single vector field. On synthetic and real spatial transcriptomics data, TP-DATE better preserves spatial structure and improves continuous 3D dynamics reconstruction.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20008v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20008v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning</title>
      <link>https://arxiv.org/abs/2609.20004v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20004v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:08:48 GMT</pubDate>
      <dc:creator>Nikita Khomich, Leopold Hermansson, Ido Hakimi</dc:creator>
      <category>模型架构</category>
      <description>Reward-based reinforcement learning for language models, exemplified by Group Relative Policy Optimization (GRPO), collapses an entire stochastic trajectory into a single scalar reward. This is clean ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Nikita Khomich, Leopold Hermansson, Ido Hakimi</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Reward-based reinforcement learning for language models, exemplified by Group Relative Policy Optimization (GRPO), collapses an entire stochastic trajectory into a single scalar reward. This is clean and scalable, but it explores and allocates reward inefficiently: a trajectory may contain many causal decisions, recovery attempts, and environment-randomness events, yet every token or action inherits one trajectory-level advantage. We study tree-based rollout construction as a compute-allocation problem for policy-gradient estimation. Our central claim is that branches should be placed not where the policy is merely uncertain, but where an additional branch most reduces uncertainty about the policy gradient per unit of compute. From a law-of-total-variance decomposition of the local policy-gradient random variable, we derive two allocation laws: new branches reduce decision uncertainty, while repeated suffix rollouts reduce continuation uncertainty. The resulting EPIG-Tree score allocates branches using the already computed rollouts. It estimates occupancy- and score-weighted value uncertainty, along with a suffix law $n_e \propto w_e \|\nabla_θ\log π(a_e|h_e)\| σ_e / \sqrt{c_e}$. Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching. In online single-turn math, tree-local credit beats flat GRPO, while branch placement is secondary to token-level credit assignment. In online multi-turn Wordle, EPIG attains the highest final win rate (0.850), overtaking flat GRPO, which saturates early at 0.790, and entropy branching as training proceeds, confirming that the gradient-estimation advantage transfers to a stateful, large-action setting.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20004v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20004v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews</title>
      <link>https://arxiv.org/abs/2609.20001v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.20001v1</guid>
      <pubDate>Thu, 17 Sep 2026 10:07:26 GMT</pubDate>
      <dc:creator>Haoshen Wang, Dongbo Che, Zeyi Xie et al.</dc:creator>
      <category>模型架构</category>
      <description>Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an eviden...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Haoshen Wang, Dongbo Che, Zeyi Xie et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.20001v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.20001v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models</title>
      <link>https://arxiv.org/abs/2609.19991v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19991v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:58:27 GMT</pubDate>
      <dc:creator>Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei et al.</dc:creator>
      <category>模型架构</category>
      <description>Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessme...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training examples and category-balanced development and test splits of 3,500 and 7,000 examples. We evaluate five open omni models under their respective input configurations using reference-blind response normalization followed by deterministic scoring. All five off-the-shelf systems score below the test split&apos;s majority-label baseline of 0.556 on synchronization verification, and obtain low scores on chain parsing and event-conditioned grounding and comprehension. Development-set perturbations reveal task-dependent sensitivity in Qwen3-Omni-30B to modality removal and changes in visual input processing, without isolating their underlying causes. Parameter-efficient temporal post-training improves Gemma4-E4B-it on several benchmark metrics. On three external image benchmarks, task metrics change modestly, including some degradations, while teacher-forcing perplexity decreases. Together, these findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19991v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19991v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning</title>
      <link>https://arxiv.org/abs/2609.19990v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19990v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:58:16 GMT</pubDate>
      <dc:creator>Shengli He, Yongchao Liang, Roumeng He et al.</dc:creator>
      <category>模型架构</category>
      <description>The high visual-token load in multimodal large language models (MLLMs) motivates training-free pruning to reduce later-layer computation, but under a fixed budget, pruning must preserve query-relevant...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shengli He, Yongchao Liang, Roumeng He et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The high visual-token load in multimodal large language models (MLLMs) motivates training-free pruning to reduce later-layer computation, but under a fixed budget, pruning must preserve query-relevant evidence while avoiding redundancy. Existing methods rank tokens, diversify selected subsets, or optimize coverage without using a shared per-visual query utility to weight both visual targets and candidate representatives. We introduce QCPruner, which makes both roles query-conditioned through bilateral utility weighting. Using keyword-matched query anchors, QCPruner fuses two cross-modal cues into utility and applies it to both visual targets and candidate representatives within visual-affinity-based coverage. The resulting nonnegative facility-location objective is monotone and submodular, retains the standard (1-1/e) greedy guarantee, and requires no model training or parameter updates. Across LLaVA-1.5, LLaVA-NeXT, LLaVA-Video, and Qwen2.5-VL, QCPruner achieves the highest average relative performance among evaluated complete-system pruning methods at every reported token budget. At 32 of 576 tokens on LLaVA-1.5-7B, it retains 96.1% of unpruned performance, versus 93.9% for the strongest evaluated baseline. At 256 of 1296 tokens on Qwen2.5-VL-7B, the corresponding values are 96.7% and 92.5%.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19990v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19990v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification</title>
      <link>https://arxiv.org/abs/2609.19985v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19985v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:57:46 GMT</pubDate>
      <dc:creator>Zhilong Zheng, Letian Tao, Yang Guan et al.</dc:creator>
      <category>模型架构</category>
      <description>Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they a...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhilong Zheng, Letian Tao, Yang Guan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive Subspace Orthogonality condition. In this paper, we introduce a purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the first order. By projecting parameter updates into the JAcobian NUll Space (JANUS), our method significantly recovers compromised historical knowledge without interfering with the underlying fine-tuning process. To overcome the local validity of the Jacobian approximation, we further propose a Multi-step Adaptive Rectification mechanism that utilizes the JANUS shift to dynamically verify the valid trust region and adjust step sizes. Coupled with our proposed ghost projection, ghost orientation comparison, and sequence-level singular value decomposition compression techniques, JANUS also achieves great temporal and spatial efficiency. Experiments demonstrate that JANUS seamlessly integrates with various fine-tuning methods, significantly mitigating the stability-plasticity dilemma by recovering historical knowledge while preserving downstream task adaptation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19985v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19985v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation</title>
      <link>https://arxiv.org/abs/2609.19974v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19974v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:48:33 GMT</pubDate>
      <dc:creator>Zitai Huang, Taiyi Su, Jian Zhu et al.</dc:creator>
      <category>模型架构</category>
      <description>Long-horizon robot manipulation requires not only stable local visuomotor control, but also continuous target tracking and reliable task progress assessment throughout execution. This challenge become...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zitai Huang, Taiyi Su, Jian Zhu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Long-horizon robot manipulation requires not only stable local visuomotor control, but also continuous target tracking and reliable task progress assessment throughout execution. This challenge becomes particularly critical when multiple objects share identical appearances and must be manipulated in a prescribed order. In such scenarios, relying solely on a limited-horizon manipulation policy is often insufficient to determine which instance should be operated on and when the task should transition to the next stage. To address this challenge, we propose MaskHarness-WAM, an instance-grounded harness for long-horizon manipulation. The proposed system connects high-level task planning with low-level manipulation policies through target masks, while leveraging visual feedback for subtask scheduling and continuous execution. Since each subtask corresponds to a different target instance, the low-level policy requires a newly established initial target mask under the updated scene at each subtask transition. The harness continuously re-observes the environment, generates, and verifies the target mask at subtask boundaries, thereby updating the instance-level spatial condition provided to the low-level policy. Furthermore, the system advances the manipulation process by switching target instances according to the verified completion status of each subtask. Experiments on a real robot platform demonstrate that MaskHarness-WAM substantially outperforms limited-horizon policies on sequential multi-object manipulation, showing its effectiveness in extending local manipulation skills to reliable long-horizon execution.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19974v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19974v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] [图像生成] Beyond the Foreground: FOV-Aware Polyp Image Synthesis via Lesion-Guided Adaptive Mucosal Context Propagation</title>
      <link>https://arxiv.org/abs/2609.19966v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19966v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:38:31 GMT</pubDate>
      <dc:creator>Tong Wang, Yuting He, Bin Ren et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Synthetic image and mask pairs can alleviate scarce colonoscopy annotations, but realistic synthesis requires preserving the supplied lesion while generating compatible mucosa. Existing foreground-gui...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Beyond the Foreground: FOV-Aware Polyp Image Synthesis via Lesion-Guided Adaptive Mucosal Context Propagation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tong Wang, Yuting He, Bin Ren et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, u-net, image synthesis, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Synthetic image and mask pairs can alleviate scarce colonoscopy annotations, but realistic synthesis requires preserving the supplied lesion while generating compatible mucosa. Existing foreground-guided methods treat all non-foreground pixels as background and rely mainly on local integration. Directly applying them to colonoscopy causes two problems: non-mucosal black regions contaminate generated tissue, and local reasoning produces inconsistent mucosal texture and illumination. We propose LAMP, the first foreground-guided framework for polyp image synthesis based on lesion-guided adaptive mucosal context propagation. LAMP explicitly separates the lesion, valid mucosa, and camera exterior using a field-of-view (FOV) mask. Lesion-to-Mucosa cross-attention extracts lesion appearance conditions for valid-mucosa locations, while FOV-constrained multidirectional Vision Receptance Weighted Key Value propagates them over legal tissue support. An adaptive gate then controls their residual fusion into the diffusion U-Net. Extensive experiments on five polyp datasets demonstrate that LAMP substantially outperforms existing methods in overall generation quality and consistently improves five downstream segmentation models. Our code will be released at https://github.com/wangtong627/LAMP.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19966v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19966v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Enhanced Knowledge Distillation for Detection Transformer via Teacher Prediction Refinement</title>
      <link>https://arxiv.org/abs/2609.19964v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19964v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:38:21 GMT</pubDate>
      <dc:creator>Yitong Xing, Yuhao Cheng, Yanping Li et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Detection Transformers (DETRs) achieve strong performance in object detection but remain challenging to deploy on edge devices due to their high computational cost. Existing DETR distillation methods ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Enhanced Knowledge Distillation for Detection Transformer via Teacher Prediction Refinement</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yitong Xing, Yuhao Cheng, Yanping Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Detection Transformers (DETRs) achieve strong performance in object detection but remain challenging to deploy on edge devices due to their high computational cost. Existing DETR distillation methods mainly focus on aligning distillation points, while largely overlooking the quality of the teacher&apos;s supervision itself. We observe that due to stage-wise non-monotonic prediction behavior in DETRs, well-localized or correctly classified predictions from earlier stages may degrade in later ones, and some negative predictions become increasingly overconfident. As a result, relying solely on the current stage&apos;s predictions yields inaccurate and inconsistent supervision. To address this issue, we propose Teacher Prediction Refinement Distillation (TPRD), a plug-and-play module that refines teacher predictions before distillation by exploiting stage-wise prediction information. TPRD improves supervision quality through Positive Prediction Correction (PPC), which corrects degraded positive predictions by restoring more accurate ones from earlier stages, ensuring reliable localization and classification signals, and Negative Prediction Suppression (NPS) suppresses the influence of overconfident negatives, preventing them from providing misleading supervision to the student. To preserve informative dark knowledge, we further introduce Maximum Dark Knowledge Preservation (MDKP), which selectively refines target-class logits while retaining non-target relations. Extensive experiments on MS COCO and PASCAL VOC demonstrate the effectiveness and robustness of the proposed method. Our code is available at https://github.com/xingyitong1/TPRD.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19964v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19964v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs</title>
      <link>https://arxiv.org/abs/2609.19961v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19961v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:31:24 GMT</pubDate>
      <dc:creator>Yuqi Ping, Tianhao Liang, Nanchi Su et al.</dc:creator>
      <category>模型架构</category>
      <description>Networked low-altitude unmanned aerial vehicles (UAVs) need reliable and adaptive decision-making capabilities to operate under uncertain observations, dynamic environments, and intermittent connectiv...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yuqi Ping, Tianhao Liang, Nanchi Su et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Networked low-altitude unmanned aerial vehicles (UAVs) need reliable and adaptive decision-making capabilities to operate under uncertain observations, dynamic environments, and intermittent connectivity, while many existing agentic systems remain limited by hallucination risks, data dependence, and weak generalization. This article investigates neuro-symbolic agentic AI (NSAAI) as a framework for combining neural grounding, symbolic reasoning, and closed-loop agentic interaction to support more reliable and adaptive UAV autonomy. We first examine its capability foundations in data efficiency, compositional generalization, continual learning, and zero-shot transfer, and then develop a reference architecture integrating task and goal management, neuro-symbolic planning, verification and metacognition, skill execution and network interaction, and shared knowledge and memory. An urban fire-inspection case implemented in LAESim illustrates how a UAV can coordinate sensing and cloud access under intermittent connectivity, reuse a verified image-delivery skill, and satisfy explicit evidence conditions before completing the mission. The results illustrate the potential of NSAAI to support reusable skills, evidence-grounded decision-making, and adaptive mission execution in networked UAV systems. We further discuss key research directions in uncertainty-aware reasoning, knowledge and skill expansion, adaptive self-monitoring, and standardized evaluation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19961v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19961v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] LapaTrack-3D: 6 DoF pre-operative shape tracking for laparoscopic surgery</title>
      <link>https://arxiv.org/abs/2609.19954v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19954v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:27:29 GMT</pubDate>
      <dc:creator>Jingwei Song, Javid Hussain Jakir, Ray Zhang et al.</dc:creator>
      <category>图像生成</category>
      <description>This work proposes a real-time 6 Degree-of-Freedom (6 DoF) tracking algorithm for monocular laparoscopic surgery. It provides alignment between intra-operative video and pre-operative data (e.g., CT)....</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">LapaTrack-3D: 6 DoF pre-operative shape tracking for laparoscopic surgery</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jingwei Song, Javid Hussain Jakir, Ray Zhang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">This work proposes a real-time 6 Degree-of-Freedom (6 DoF) tracking algorithm for monocular laparoscopic surgery. It provides alignment between intra-operative video and pre-operative data (e.g., CT). The 6 DoF tracking offers a solution for accurately locating the internal anatomy of the target organ despite the lack of tactile feedback and transparency. The ORB-SLAM2 framework is adopted and modified for prior-based 3D tracking with four major modifications. First, the primitive 3D shape is used for fast initialization of the ORB-SLAM2 monocular mode. Second, a pseudo-segmentation strategy is employed to separate the target organ from the background for tracking. Third, the 3D shape is incorporated as a geometric prior in its pose graph optimization. Fourth, the Multi-Scale Retinex with Chromaticity Preservation (MSRCP) algorithm is leveraged and modified for image enhancement in challenging illumination scenarios. In-vivo and ex-vivo experiments validate that LapaTrack-3D provides robust 3D tracking and effectively handles typical challenges such as poor illumination, fast motion, out-of-field-of-view scenarios, partial visibility, and ``organ-background&apos;&apos; relative motion. LapaTrack-3D achieves a processing rate of 13 Hz for 1280*720 pixel video.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19954v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19954v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation</title>
      <link>https://arxiv.org/abs/2609.19944v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19944v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:17:15 GMT</pubDate>
      <dc:creator>Yudai Nakada, Yuichiro Nishiura, Jin Michael Splichal</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Large language models (LLMs) have been applied to causal discovery, but candidate-graph generation rarely treats premature omission of potentially relevant causal relations as an explicit design objec...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yudai Nakada, Yuichiro Nishiura, Jin Michael Splichal</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Large language models (LLMs) have been applied to causal discovery, but candidate-graph generation rarely treats premature omission of potentially relevant causal relations as an explicit design objective. We propose MaSCoD, a multi-agent framework that organizes candidate third variables and local structural patterns before direct-edge judgment. We evaluate MaSCoD on Auto-MPG, DWD, and Sachs using GPT-5.4 as the primary backbone and GPT-4o for replication. MaSCoD exhibits a dataset- and backbone-dependent retention-selectivity profile rather than uniform superiority. Across all six dataset-backbone settings, Full, which supplies structural hypotheses before direct-edge judgment, achieved higher mean Recall and F1 than No Phase 1, which instead constructs them within the judgment procedure, while also increasing false-positive rates. Additional reference-edge retention over all evaluated baselines was observed on DWD with GPT-5.4 and on Sachs with GPT-4o, rather than uniformly across settings. Partial ablations showed that supplying both information components did not always outperform supplying only one. For GPT-5.4, stage-wise analysis showed that the Full-No Phase 1 retention gap was already present after direct-edge judgment, while reconciliation introduced additional reference-edge loss for Full on Sachs. These findings support structural pre-organization as an explicit design and evaluation target for omission control and motivate evaluating context construction jointly with its utilization in judgment.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19944v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19944v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute</title>
      <link>https://arxiv.org/abs/2609.19942v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19942v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:17:00 GMT</pubDate>
      <dc:creator>Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak et al.</dc:creator>
      <category>扩散模型</category>
      <description>In extractive document question answering whose questions were generated from the passages that contain their answers -- so that retrieval recovers 92-99.8% of what any mode combination could reach, w...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">In extractive document question answering whose questions were generated from the passages that contain their answers -- so that retrieval recovers 92-99.8% of what any mode combination could reach, whatever its absolute accuracy -- confidence-driven mechanisms have little to gain. Fine-tuning an open language model on a specialized domain corpus yields a model whose own confidence is a tempting control signal: it could decide which queries warrant further adaptation, and which answers to trust. We evaluate both uses under criteria fixed before the runs were executed, across four 7-9B model families whose adaptation moved closed-book F1 by at most +0.03, and both fail: a distillation trigger on all four families, under its pre-specified three-step transfer budget, and a routing-and-abstention policy in its single-model pilot. Retrieval alone recovers 92-99.8% of best-case combined accuracy under every correctness criterion we test, leaving routers no meaningful gain. The sequence-likelihood signal is insufficient relative to that mode -- area under the receiver operating characteristic curve 0.65-0.81 under the registered criterion -- before adaptation as well as after, unchanged by scalar recalibration and not consistently improved by token-level temperature rescaling. And the finer diagnostics depend on the correctness criterion and on answer length; on the three adapted combinations where we could test it, selector ablations show no statistically detectable downstream benefit from the confidence term on any seed; on Gemma, removing it changes the selector from failing to passing both registered criteria. The usable product is a set of pre-specified negatives with their dependencies made explicit.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19942v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19942v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Stringological sequence prediction III: layered ziplines and a tradeoff between efficiency and expressivity</title>
      <link>https://arxiv.org/abs/2609.19940v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19940v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:15:29 GMT</pubDate>
      <dc:creator>Vanessa Kosoy</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>In previous papers, we began the study of sequence prediction algorithms adapted to stringological word complexity measures. In particular, we defined a complexity measure called Arithmetic Repetition...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Stringological sequence prediction III: layered ziplines and a tradeoff between efficiency and expressivity</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Vanessa Kosoy</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">In previous papers, we began the study of sequence prediction algorithms adapted to stringological word complexity measures. In particular, we defined a complexity measure called Arithmetic Repetition Complexity (ARC) which admits a polynomial-time prediction algorithm with a mistake bound quasilinear in the complexity. Here, we show a weaker complexity measure related to ARC that admits an especially efficient prediction algorithm: an algorithm that runs in quasilinear time and polylog space for appropriate highly-structured sequences. The complexity measure is defined via a restricted class of &quot;zipline programs&quot; (a variant of straight-line programs), which we call layered. We thus get a less expressive measure with a more efficient algorithm (compared to our results for ARC), demonstrating a possible tradeoff.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19940v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19940v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models</title>
      <link>https://arxiv.org/abs/2609.19934v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19934v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:09:43 GMT</pubDate>
      <dc:creator>Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh Dung</dc:creator>
      <category>模型架构</category>
      <description>Depth-recurrent language models iteratively apply a small layer stack, decoupling per-token compute from distinct parameter count. To determine whether such a model genuinely utilizes its depth, both ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh Dung</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Depth-recurrent language models iteratively apply a small layer stack, decoupling per-token compute from distinct parameter count. To determine whether such a model genuinely utilizes its depth, both recurrence and layer-pruning literatures rely on a shared evaluation: truncating depth at inference time, plotting quality against retained depth fraction, and reading off the slope. While cheap and training-free, this metric suffers from an unexamined flaw: it extracts a single scalar from an intervention that alters multiple model properties simultaneously. Depth truncation concurrently reduces the number of block applications, decreases the volume of distinct computation performed, and pushes the readout head onto an out-of-distribution residual stream. The observed slope conflates all three factors, yet is conventionally interpreted as reflecting solely the second.   We propose the Depth Control Protocol (DCP), a diagnostic suite that disentangles these three quantities. DCP comprises three positive controls that isolate each factor while varying the others, a negative control applying the identical interventions to dense transformers to ensure the effect is not an artifact of the measurement protocol, and a controlled training intervention to verify causality. The linchpin control, running the full budget of block applications while executing only a single distinct iteration, is strictly realizable only in depth-wise weight-sharing architectures, since in a dense network repeating a layer yields an entirely different model rather than the same model in an alternative configuration.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19934v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19934v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] DirtyMoCap: Robust Motion Capture from Unconstrained Markers</title>
      <link>https://arxiv.org/abs/2609.19927v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19927v1</guid>
      <pubDate>Thu, 17 Sep 2026 09:05:42 GMT</pubDate>
      <dc:creator>Long Wang, Shuting Zhao, Shen Yan et al.</dc:creator>
      <category>模型架构</category>
      <description>Optical motion capture delivers high-fidelity human motion, but its reliance on strict marker layouts and clean trajectories severely limits its real-world applicability. In practice, tracking systems...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">DirtyMoCap: Robust Motion Capture from Unconstrained Markers</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Long Wang, Shuting Zhao, Shen Yan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Optical motion capture delivers high-fidelity human motion, but its reliance on strict marker layouts and clean trajectories severely limits its real-world applicability. In practice, tracking systems frequently output unconstrained markers: sparse, noisy, and unordered point clouds with unknown or varying configurations. To bridge the gap between corrupted raw markers and parametric human models, we introduce DirtyMoCap, a robust, marker-layout-free framework. Our core insight is to map unordered marker observations to a fixed set of &quot;proxy anchors&quot; comprising skeletal joints and body surface points, which serve as a stable intermediate representation. We first initialize and track these anchors over long sequences using a recurrent sliding-window architecture. Then, a custom differentiable Gauss-Newton solver fits the SMPL-H model to the tracked anchors to recover full-body pose, translation, and shape. By explicitly deriving geometric residuals, our solver learns adaptive observation confidence, smoothness, and prior weights end-to-end, adapting dynamically to the reliability of the input data. Extensive experiments on diverse, noisy marker configurations demonstrate that DirtyMoCap successfully generalizes across arbitrary layouts using only a single trained model. It consistently outperforms state-of-the-art configuration-specific baselines in both joint and vertex reconstruction accuracy, while our custom CUDA solver achieves up to a 100x speedup over standard PyTorch implementations. We further apply DirtyMoCap to heterogeneous raw optical MoCap recordings of traditional Chinese martial arts, yielding a Kung Fu motion dataset of temporally coherent SMPL-H reconstructions. Code and data are available at https://wanglongzju.github.io/DirtyMoCap-Project-Page.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19927v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19927v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Amortizing Physics-Informed Neural Solvers via Graph Hypernetworks</title>
      <link>https://arxiv.org/abs/2609.19915v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19915v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:58:01 GMT</pubDate>
      <dc:creator>Cheng Jing, Abhishek Verma, Kallol Bera et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Amortizing physics-informed neural networks (PINNs) across related PDEs requires describing each equation to a reusable solver. Coefficient vectors encode numerical parameters in predefined slots, lea...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Amortizing Physics-Informed Neural Solvers via Graph Hypernetworks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Cheng Jing, Abhishek Verma, Kallol Bera et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Amortizing physics-informed neural networks (PINNs) across related PDEs requires describing each equation to a reusable solver. Coefficient vectors encode numerical parameters in predefined slots, leaving operator and cross-field assignments implicit. We make these relationships explicit in an operator graph, with nodes for fields, derivatives, terms, and residuals and coefficients retained as term attributes. A graph hypernetwork generates diagonal codes that initialize a meta-trained factorized PINN for each target equation. Meta-training and target-specific adaptation use governing equations and prescribed conditions without solution labels. We compare coefficient-vector, DeepSets-based term-set, and graph conditioning by solution accuracy within a fixed adaptation budget. In scalar convection-diffusion-reaction problems, both term-based descriptors improve high-reaction accuracy, with similar performance. In two-field Fisher-KPP, meta-training sees uncoupled and one-way systems; after 3,000 adaptation steps on unseen two-way coupling, the graph&apos;s mean final error is 35.7% below the term set and 67.7% below the coefficient vector. In a fixed-structure capacitively coupled plasma model, the coefficient vector performs best. These results support extending coefficient conditioning with explicit equation relationships for physics-based solver adaptation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19915v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19915v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks</title>
      <link>https://arxiv.org/abs/2609.19913v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19913v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:53:47 GMT</pubDate>
      <dc:creator>Omran Berjawi, Giuseppe Fenza, Rida Khatoun et al.</dc:creator>
      <category>模型架构</category>
      <description>The study of opinion dynamics in social networks is one of the key challenges in computational social science with direct relevance to understanding political polarization, misinformation, and health ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Omran Berjawi, Giuseppe Fenza, Rida Khatoun et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> mae, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The study of opinion dynamics in social networks is one of the key challenges in computational social science with direct relevance to understanding political polarization, misinformation, and health responses. Current approaches focus on simplified mathematical models that ignore linguistic and contextual factors related to belief updates or use Large Language Model (LLM)-based simulations that have not been validated against real data. We present a framework based on the concept of a digital twin to simulate opinion dynamics in social networks. The approach fills the gap by cloning a real-world Twitter network, assigns a set of attributes for agents (such as persona, emotions, centrality, stubbornness, and influence), and employs Mistral-7B to perform opinion update based on memory and social exposure. To evaluate the proposed approach, we validate it against two real Twitter datasets (COVID-19 discourse and U.S elections 2020). The results show that the capability of the proposed framework reproduces opinion trajectories and reduces individual prediction error by more than 50% compared to the best-performing classical baseline (Mistral-7B achieves Mean Absolute Error (MAE) = 0.150 and 0.121 on the COVID-19 and US Election 2020 datasets, respectively). We observe similar improvements in structural alignment (Delta_r = 0.120 and 0.180) and polarization dynamics (Delta_Var = 0.106 and 0.115) on the two datasets, respectively. Additionally, the ablation studies confirm that agent attributes, memory, and social exposure all contribute to the framework&apos;s predictive fidelity in reproducing opinion trajectories, with agent attributes being the most critical contributor. Overall, our results demonstrate that grounding Mistral-7B within empirically cloned interaction networks produces a realistic simulation framework capable of reproducing complex social dynamics.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19913v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19913v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding</title>
      <link>https://arxiv.org/abs/2609.19911v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19911v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:53:13 GMT</pubDate>
      <dc:creator>Shuai Zhang, Hongye Hou, Qinghe Liu et al.</dc:creator>
      <category>图像生成</category>
      <description>3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rel...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shuai Zhang, Hongye Hou, Qinghe Liu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit semantic and geometric structures hidden in billion-scale urban point clouds. We reformulate city-scale 3D grounding as structured constraint reasoning, where description semantics are organized into computable cross-modal constraints over open-vocabulary 3D entities, attributes, and spatial relations. We present CitySTAR, a training-free framework for reasoning-driven urban 3D grounding. CitySTAR lifts raw billion-scale urban point clouds into a query-ready scene graph of open-vocabulary 3D instances, with CodeLLM-driven tools supplying multimodal evidence for node attributes and 3D spatial relations. It then models target-context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation. Finally, a Reflective Cross-modal Grounding module integrates topology consistency and candidate-centered 2D visual evidence to make decisions over a metric-aware 3D context graph. To further support this setting, we introduce CitySTAR-3D, an enhanced benchmark that improves semantic coverage, instance completeness, bounding-box fidelity, and spatial-relation complexity in city-scale 3D grounding. Extensive experiments show that CitySTAR consistently improves open-world urban 3D grounding while maintaining strong interpretability and generalization.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19911v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19911v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] GS-PI: An Optimization-Decoupled Appearance Decomposition Approach for Generating PBR Gaussian Assets</title>
      <link>https://arxiv.org/abs/2609.19907v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19907v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:51:18 GMT</pubDate>
      <dc:creator>Jieting Xu, Rengan Xie, Zijian Huang et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Gaussian Splatting (GS) excels at novel-view synthesis but encodes baked-in radiance, tightly entangling illumination with geometry and preventing seamless integration into physically based rendering ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">GS-PI: An Optimization-Decoupled Appearance Decomposition Approach for Generating PBR Gaussian Assets</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jieting Xu, Rengan Xie, Zijian Huang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Gaussian Splatting (GS) excels at novel-view synthesis but encodes baked-in radiance, tightly entangling illumination with geometry and preventing seamless integration into physically based rendering (PBR) pipelines. Existing inverse-rendering methods attempt to disentangle materials via joint optimization, but often suffer from competing objectives that cause severe ambiguities and residual lighting artifacts. To overcome this, we present GS-PI, a novel optimization-decoupled framework that casts PBR material generation as a geometry-conditioned diffusion process on 3D point clouds. By operating directly in the 3D domain, our method inherently guarantees multi-view consistency, sidestepping the severe pixel correspondence issues that challenge 2D diffusion approaches. We introduce a multi-scale cross-view conditioning mechanism that integrates three complementary components: a global semantic prior, source-anchored photometric cues, and an absolute spatial learned view-direction conditioning signal. This design efficiently compresses complex multi-view evidence, mitigating cross-view projection misalignment and successfully preventing specular highlights from baking into intrinsic colors. By extracting a point cloud from a pre-trained Gaussian model, predicting PBR attributes via conditional diffusion, and distilling them back through differentiable rasterisation, we yield a fully relightable PBR-GS asset. GS-PI outperforms recent inverse-rendering baselines while replacing per-scene joint illumination/BRDF optimization with a learned diffusion pass followed by a short target-driven distillation, without requiring proxy meshes.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19907v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19907v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Self-Replicating Neural Cellular Automata: Quantifying Emergent Phenotypic and Genotypic Diversity in an OpenEnded Substrate</title>
      <link>https://arxiv.org/abs/2609.19902v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19902v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:46:26 GMT</pubDate>
      <dc:creator>Sanyam Jain, Felix Simon Reimers, Stefano Nichele</dc:creator>
      <category>图像生成</category>
      <description>We study an in-silico substrate in which every pixel of a two-channel cellular-automata grid carries a tiny neural network (an agent) that senses its Moore neighborhood. A cell persists only by self-r...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Self-Replicating Neural Cellular Automata: Quantifying Emergent Phenotypic and Genotypic Diversity in an OpenEnded Substrate</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sanyam Jain, Felix Simon Reimers, Stefano Nichele</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We study an in-silico substrate in which every pixel of a two-channel cellular-automata grid carries a tiny neural network (an agent) that senses its Moore neighborhood. A cell persists only by self-replication: a living neighbor is cloned and its weights are mutated by a uniform perturbation, so that phenotype (cell state) is driven entirely by genotype (network weights). From a handful of seeded founders the system grows into a spatially organized ecosystem of coexisting, competing and dominating species. Our main contribution is a battery of coarse-grained diversity metrics that make such growth measurable at two scales: four phenotypic tools based on cellular-type frequency, entropy and cell variance, and two genotypic tools that colour each agent by a hash of its full weight vector versus a sparse random-weight probe. Across a five-fold sweep of 1680 small runs and 24 long (1000-generation, 200 x 200) runs, the substrate is persistent and self-maintaining in 20 of the 24 long configurations and exposes a clear phenotype-genotype diversity trade-off: raising phenotypic diversity collapses genotypic diversity and vice versa. Full-genome hash colouring further reveals lineage structure that a random-weight probe systematically misses. Code, data and animations are released as supplementary material.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19902v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19902v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Trigger Timing, Deadline Readiness, and Event-Aligned Accounting for Dynamic Ad Insertion</title>
      <link>https://arxiv.org/abs/2609.19899v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19899v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:44:52 GMT</pubDate>
      <dc:creator>Prashant Chaudhary, Kapil Khandelwal</dc:creator>
      <category>模型架构</category>
      <description>Dynamic ad insertion comparisons can conflate trigger, reach, readiness, playback, billability and measurement even when the accounting is arithmetically correct. We separate these events with an obse...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Trigger Timing, Deadline Readiness, and Event-Aligned Accounting for Dynamic Ad Insertion</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Prashant Chaudhary, Kapil Khandelwal</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Dynamic ad insertion comparisons can conflate trigger, reach, readiness, playback, billability and measurement even when the accounting is arithmetically correct. We separate these events with an observed-event ledger, a candidate-invariant reference deadline and pod-level contribution accounting. The deadline rule is fixed before candidate assignment and tests whether an admissible transition state remains valid, not whether preparation merely finished earlier. A restricted monotone-playback representation states when media-position summaries suffice; a counterexample shows why they fail over a wider path class. An offline synthetic study exercises the definitions over nine short-lifetime conditions informative for the readiness comparison and nine long-lifetime conditions serving as analytic controls. Across 45,000 shared scripts, two trigger policies share imposed playback paths, latency draws and hypothetical value and cost coefficients. Playhead summaries substantially misclassify reach events in the nonmonotone mixtures, yet neither shortcut reverses the contribution contrast in this grid, because some errors cancel under the shared design. Replacing validity at the deadline with completion by the deadline reverses the comparison in three of the nine informative conditions, all at one of the three latency settings. Scoring readiness at actual viewer arrival rather than at the reference deadline shifts pause-path readiness but changes no contribution sign. These outcomes are consequences of the event definitions applied to established misclassification mechanisms. Conservative bounds retain uncertainty when records are missing, and the artifact records code, seeds, event histories and checking procedures. The evidence is synthetic, uses no commercial telemetry, and ranks neither server-side nor server-guided insertion.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19899v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19899v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Online Adaptive Kernel Mixing for Gaussian Process Decision Making</title>
      <link>https://arxiv.org/abs/2609.19891v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19891v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:38:10 GMT</pubDate>
      <dc:creator>Kavin Aravindan, Mani Tej Sriram, Gautam Dasarathy et al.</dc:creator>
      <category>模型架构</category>
      <description>Gaussian Processes (GPs) are widely used as surrogates for black-box functions in sequential decision-making problems such as Bayesian optimization (BO), level set estimation (LSE), and Bayesian activ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Online Adaptive Kernel Mixing for Gaussian Process Decision Making</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kavin Aravindan, Mani Tej Sriram, Gautam Dasarathy et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Gaussian Processes (GPs) are widely used as surrogates for black-box functions in sequential decision-making problems such as Bayesian optimization (BO), level set estimation (LSE), and Bayesian active learning (BAL). GP performance critically depends on kernels, and standard kernels can lead to suboptimal decisions under misspecification. To address this, we introduce HACK GPs (Hedge Adaptive Cumulative Kernels), a method that views kernel selection as an online learning with expert advice problem. HACK treats each candidate kernel as a GP &quot;expert&quot; and updates a distribution over experts online using AdaHedge, based on a loss received as a proxy for their ability to fit the function and align with the task objective. We provide two variants of HACK: (i) Mixture of Gaussians (MoG) and (ii) categorical sampling. We establish general guarantees showing that, under a loss-gap condition, the weight concentrates on the best kernel and the resulting acquisition function is close to that of the best expert. Empirically, we observe robust performance across BO, LSE, and BAL compared to standard kernels such as Squared Exponential and Matern-5/2, as well as simple ensemble baselines.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19891v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19891v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces</title>
      <link>https://arxiv.org/abs/2609.19883v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19883v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:31:18 GMT</pubDate>
      <dc:creator>Pyrros Koussios, Benjamin Jäger, John Hua Yao et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a c...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Pyrros Koussios, Benjamin Jäger, John Hua Yao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained, and scalable benchmark for evaluating LLM reasoning over dynamic state spaces using Petri nets, a mature formalism for modeling real-world concurrent and distributed systems. PetriBench organizes reasoning into four task families varying by scope and temporal horizon, with Easy, Medium, and Hard levels generated by increasing structural complexity and evaluated against exact ground truth. Across a diverse set of proprietary and open-weight models, accuracy decreases consistently with difficulty, while harder instances expose increasingly distinct task-specific capability profiles. Additional analyses show that test-time compute improves performance but interacts differently with different reasoning tasks, and that procedural generation yields smooth scaling with structural complexity. Together, these results show that PetriBench provides a unified and extensible setting for probing the strengths, limits, and scaling behavior of LLM reasoning.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19883v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19883v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [扩散模型] [图像生成] Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning</title>
      <link>https://arxiv.org/abs/2609.19878v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19878v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:28:18 GMT</pubDate>
      <dc:creator>Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang et al.</dc:creator>
      <category>多模态生成</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a sing...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, vision-language model, latent diffusion, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a single sequence, leaving the model to bridge representational differences as it reasons across modalities. We introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that brings these thoughts into a shared latent space for reasoning. A unified encoder maps teacher reasoning steps from different modalities into shared thought tokens, trained to preserve the information needed for later reasoning steps and the final answer or action. Because the same context can support multiple valid next steps, we use diffusion to predict the next block of thought tokens from the input and preceding blocks. Jointly training the encoder and diffusion reasoner with shared model weights encourages thought tokens to be both useful for the task and predictable from the available context. At inference, the model generates these tokens without teacher observations. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest evaluated baselines of 7.3% on visual reasoning tasks and 6.1% on robot manipulation tasks.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19878v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19878v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection</title>
      <link>https://arxiv.org/abs/2609.19873v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19873v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:26:04 GMT</pubDate>
      <dc:creator>Omran Berjawi, Walid fahs, Rida Khatoun</dc:creator>
      <category>模型架构</category>
      <description>Email spam and phishing attacks remain a critical security threat. Adversaries increasingly exploit large language models to craft contextually convincing malicious messages, and existing spam detecti...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Omran Berjawi, Walid fahs, Rida Khatoun</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Email spam and phishing attacks remain a critical security threat. Adversaries increasingly exploit large language models to craft contextually convincing malicious messages, and existing spam detection systems often struggle to keep pace. Generalization across diverse and evolving attack scenarios is limited, which reduces effectiveness once these systems are deployed in practice. This paper introduces Adaptive Uncertainty-Routed Analysis (AURA), a multimodal email threat detection system that analyzes both the content of an email and its embedded URLs. AURA is built around two layers: the first quantifies prediction uncertainty from a URL classifier, and only ambiguous messages are escalated to a fine-tuned transformer encoder for semantic analysis. The system is evaluated on eight heterogeneous training corpora together with two held-out real-world corpora spanning a decade of adversarial campaigns. AURA reaches a macro F1-score of 0.9858 in-distribution, and on NazPhish-Eval and GuenterTrap-Eval it maintains 0.9502 and 0.9436, respectively, which is evidence of robust generalization under genuine distribution shift.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19873v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19873v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] PART: Learning 3D Part Assembly and Retrieval with Transformers</title>
      <link>https://arxiv.org/abs/2609.19872v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19872v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:23:25 GMT</pubDate>
      <dc:creator>Ruchao Bao, Wenzheng Wu, Chucheng Xiang et al.</dc:creator>
      <category>模型架构</category>
      <description>3D assembly is fundamental to modern manufacturing and digital content creation. In this paper, we present PART, a unified transformer-based framework for 3D part retrieval and assembly: given a targe...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PART: Learning 3D Part Assembly and Retrieval with Transformers</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ruchao Bao, Wenzheng Wu, Chucheng Xiang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">3D assembly is fundamental to modern manufacturing and digital content creation. In this paper, we present PART, a unified transformer-based framework for 3D part retrieval and assembly: given a target shape and a part library, PART automatically selects the appropriate parts and predicts their 6-DoF poses to reconstruct the target. While prior work has achieved impressive progress on assembling a pre-defined set of parts, this more practical retrieval-based setting remains largely unexplored. The task faces three key challenges: (i) a combinatorially explosive search space that grows exponentially with library size; (ii) variable-length outputs, as different targets require different numbers of parts; and (iii) continuous 6-DoF pose estimation for part assembly. To address these, we formulate retrieval and assembly as a set prediction problem and design a novel transformer-based framework that retrieves parts and regresses their poses with variable-length output. Additionally, we exploit the duality between part pose estimation and target segmentation through joint training and a novel segmentation-enhanced optimization module. Finally, We curate a large-scale dataset of 80K+ shapes, and the results show that PART generalizes to scene layouts, image targets, and real-world scans. Project Page: https://iambrc.github.io/PART-project-page/.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19872v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19872v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Physical knowledge on historical data matters more than enforcing physical constraints on the forecast</title>
      <link>https://arxiv.org/abs/2609.19871v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19871v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:23:08 GMT</pubDate>
      <dc:creator>Etienne Lehembre, Pascal Audigane, Vincent Nguyen et al.</dc:creator>
      <category>模型架构</category>
      <description>Time series forecasting has seen signicant advancements with the emergence of new deep learning models. However, forecasting time series in applications involving physical processes remains a major ch...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Physical knowledge on historical data matters more than enforcing physical constraints on the forecast</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Etienne Lehembre, Pascal Audigane, Vincent Nguyen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Time series forecasting has seen signicant advancements with the emergence of new deep learning models. However, forecasting time series in applications involving physical processes remains a major challenge. Despite the apparition of Physics Informed Neural Networks (PINN), recent models do not estimate unobservable intermediate physical variables, which are important for domain experts to understand the target behavior. To this end, we propose a Physics Informed Recurrent Neural Network (PIRNN) which predicts, along the target, unobservable variables on both historic data and forecast target. This approach enhances the model robustness and results interpretation using domain knowledge. Our method is easily adaptable to any physical model using several equations, each having its own set of unobservable variables, to describe it-self. As a case study, we incorporate physical equations used for groundwater levels predictions by the physical model called Gardenia. This model uses transfers equations between reservoirs, optimized with data assimilation, to simulate the evolution of groundwater levels. Evaluation includes several well known neural network models and the Gardenia model compared on twelve real world datasets. In addition, we study the impact of each component through an ablation study. Our model outperforms other models on ve out of the twelve datasets and our ablation study underlines the importance of having a physical background in our time series forecasting task. Finally, the coherence of the physical variables predicted by our neural network is assessed by a domain expert.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19871v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19871v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] [图像生成] Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference</title>
      <link>https://arxiv.org/abs/2609.19868v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19868v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:21:14 GMT</pubDate>
      <dc:creator>Leonid Sinev, Ilya Koziev, Vladislav Leshchuk</dc:creator>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Leonid Sinev, Ilya Koziev, Vladislav Leshchuk</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, diffusion model, lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead due to the inability to reuse Key-Value (KV) cache and from incoherent generation arising from learning dependencies over an intractable space of token combinations. We introduce Zarya, a family of hybrid language models that jointly optimizes an autoregressive (AR) objective and a masked-diffusion objective within a single architecture. Zarya structures training data into variable-size slots and employs a curriculum that gradually increases slot granularity, enabling a smooth transition from fine-grained AR learning to coarse-grained diffusion learning. At inference, Zarya provides two distinct decoding paradigms through a unified interface: (i) MDM sampling with first-hitting denoising, and (ii) slotted speculative decoding that interleaves inter-slot diffusion-based selection with intra-slot autoregressive infilling, achieving full KV cache reuse. The training and inference regimes are fully decoupled, allowing a model trained with any configuration to be deployed in either mode. Extensive configurability --- including grouped noise patterns (Prefix Completion, Fill-In-the-Prefix, Fill-In-the-Middle), ordered sampling schedules, and noise-level permutation strategies --- enables flexible research exploration. We release Zarya models publicly in sizes 0.6B, 1.7B, and 4B, demonstrating performance on standard benchmarks while offering a principled integration of autoregressive and diffusion paradigms.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19868v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19868v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] Socialized UAV Cross-Task Learning: Towards Cross-Granularity Collaboration through Hierarchical Interaction</title>
      <link>https://arxiv.org/abs/2609.19867v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19867v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:20:51 GMT</pubDate>
      <dc:creator>Xinjie Yao, Ruipu Zhao, Yunqi Zhu et al.</dc:creator>
      <category>扩散模型</category>
      <description>Joint learning across heterogeneous tasks is often treated as task coupling through feature sharing, distillation, or auxiliary supervision. However, in cross-task learning, mismatched representationa...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Socialized UAV Cross-Task Learning: Towards Cross-Granularity Collaboration through Hierarchical Interaction</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xinjie Yao, Ruipu Zhao, Yunqi Zhu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Joint learning across heterogeneous tasks is often treated as task coupling through feature sharing, distillation, or auxiliary supervision. However, in cross-task learning, mismatched representational and supervisory granularities make such coupling prone to interference, teacher bias, or unidirectional collapse. We argue that cross-granularity learning is fundamentally a problem of hierarchical interaction regulation rather than simple task coupling. This issue is particularly evident in UAV perception, where visual shifts and detection--segmentation objectives naturally form coarse- and fine-grained knowledge sources. To systematically study this problem, we introduce CrossUAV, a UAV benchmark for joint object detection and instance segmentation that provides a unified evaluation platform for cross-granularity task collaboration. To address these challenges, we propose Cross-Granularity Socialized Collaboration (CGSC), a progressive and adaptive framework that regulates when, where, and how tasks exchange information across network hierarchies. CGSC progressively activates cross-task interactions and adaptively adjusts the strength according to task contribution, suppressing harmful interference while exploiting complementary coarse- and fine-grained structures. Extensive experiments demonstrate consistent improvements on both tasks, validating hierarchical dynamic interaction as an effective mechanism for cross-granularity collaboration.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19867v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19867v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Reproducibility is not construct validity: LLM measurement of institutionally situated communication</title>
      <link>https://arxiv.org/abs/2609.19866v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19866v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:19:58 GMT</pubDate>
      <dc:creator>Veronika Batzdorfer, Carlo Romano Marcello Alessandro Santagiustina</dc:creator>
      <category>模型架构</category>
      <description>High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Com...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Reproducibility is not construct validity: LLM measurement of institutionally situated communication</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Veronika Batzdorfer, Carlo Romano Marcello Alessandro Santagiustina</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission&apos;s AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations &gt; 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran&apos;s I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19866v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19866v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Pretrained Medical Representations for the Practical Screening of Drug Repositioning Candidates</title>
      <link>https://arxiv.org/abs/2609.19865v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19865v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:19:34 GMT</pubDate>
      <dc:creator>Yuhei Fujioka, Daitaro Misawa, Shingo Fukuma</dc:creator>
      <category>图像生成</category>
      <description>Representation learning from medical code sequences in electronic health records and medical claims data has been successful in various clinical applications, such as those regarding disease predictio...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Pretrained Medical Representations for the Practical Screening of Drug Repositioning Candidates</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yuhei Fujioka, Daitaro Misawa, Shingo Fukuma</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Representation learning from medical code sequences in electronic health records and medical claims data has been successful in various clinical applications, such as those regarding disease prediction. However, significant challenges remain in extending this approach to the discovery of scientific hypotheses. One reason is that many existing BERT-based models fail to adequately capture the hierarchical structure of medical codes and the complex interactions between diagnoses and treatments. To address these limitations, we propose a new unified pre-training framework that explicitly integrates hierarchical sub-token aggregation, partial masking, and cross-reference mechanisms. The proposed model consistently outperformed existing methods on both pre-training objectives and downstream clinical event prediction tasks, including the onset of dementia and hospitalization. We also conducted an in silico drug repositioning case study targeting Alzheimer&apos;s disease. In the hypothesis generation step, our approach successfully rediscovered known promising drugs in a data-driven manner without relying on such external knowledge sources as the literature. Subsequently, in the hypothesis prioritization step, we introduced a Task-Adaptive Representation Approach to alleviate the over-encoding of historical prescription information within diagnostic vectors, enabling the robust prioritization of generated hypotheses. This study establishes an exploratory screening workflow for hypothesis generation and prioritization based on observational associations. Importantly, this framework is not intended to provide causal evidence, but rather to identify promising candidates for subsequent rigorous causal inference. Overall, this study demonstrates that domain-informed representation learning combined with task-adaptive representation control can enable a practical hypothesis discovery workflow.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19865v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19865v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Feeling Terrain Before Crossing: World Models for Off-Road Navigation</title>
      <link>https://arxiv.org/abs/2609.19863v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19863v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:19:06 GMT</pubDate>
      <dc:creator>E-In Son, Dong-Wook Kim, Ji-Hoon Hwang et al.</dc:creator>
      <category>模型架构</category>
      <description>Navigation world models plan by foresight, predicting the future that each candidate action sequence produces and selecting the best, rather than mapping observations to actions directly. Unlike urban...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Feeling Terrain Before Crossing: World Models for Off-Road Navigation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> E-In Son, Dong-Wook Kim, Ji-Hoon Hwang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Navigation world models plan by foresight, predicting the future that each candidate action sequence produces and selecting the best, rather than mapping observations to actions directly. Unlike urban settings where a predicted scene is a sufficient proxy, off-road navigation hinges on the robot--terrain interaction, so the prediction must cover not only what the camera will see but what the robot will feel. However, existing scene-focused models do not predict how much the robot will slip, tilt or shake along a planned trajectory. Proprioception captures these dynamics directly and, when used as input, improves the prediction of the physical future. We present Feel-WM, the first off-road navigation world model that conditions on proprioception and predicts what the robot will feel alongside what the camera will see. The physical future takes the form of a future proprioceptive state and a failure risk, both learned from the robot&apos;s own experience without human labels. The planner rolls out the physical future alongside the scene and weighs the predicted failure risk against goal similarity in a separable score. Experiments on real off-road data and in simulation demonstrate that Feel-WM outperforms visual-only navigation world models in open-loop planning and closed-loop rough-terrain navigation across wheeled and legged platforms. Deployed on a Husky on mountain trails, Feel-WM plans onboard, predicts rough ground ahead and steers around it, completing courses that an end-to-end policy fails.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19863v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19863v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Expected Hypervolume Maximization for Multiobjective Optimization under Uncertainties</title>
      <link>https://arxiv.org/abs/2609.19858v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19858v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:10:15 GMT</pubDate>
      <dc:creator>Victor Trappler</dc:creator>
      <category>模型架构</category>
      <description>The problem of multiobjective optimization under uncertainties is often approached by taking the expectation of each objective. In this work, we propose instead to formulate this as a Bayesian decisio...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Expected Hypervolume Maximization for Multiobjective Optimization under Uncertainties</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Victor Trappler</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The problem of multiobjective optimization under uncertainties is often approached by taking the expectation of each objective. In this work, we propose instead to formulate this as a Bayesian decision problem and to rely on the expected value of the hypervolume, which is to be maximized with respect to a finite set of input points. We show that this can be performed using methods based on gradients in a stochastic optimization framework, provided that care is taken with respect to dominated points. Moreover, in the absence of readily available differentiable code, we propose to use Gaussian Processes as differentiable surrogate models, in order to perform the optimization. An additional contribution in this work are some active learning strategies, through acquisition functions which helps construct a surrogate model well-designed for the multiobjective optimization problem at stake. These strategies are compared on simple analytical problems to assess their performances.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19858v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19858v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] [图像生成] PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance</title>
      <link>https://arxiv.org/abs/2609.19853v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19853v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:05:08 GMT</pubDate>
      <dc:creator>Bing Duan, Qiang Guo, Linpu Li et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text set...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Bing Duan, Qiang Guo, Linpu Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, diffusion model, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the plan: the screenplay evidence, the characters, props and locations it needs, where each subject stands, and what the camera does. A value is written once at the level it belongs to (script, scene, shot or panel) and inherited below it. A compiler turns the result into both the prompt sent to the diffusion model and a 3D scene built in metres, and a camera solver places the camera so that the declared framing is the framing built. Where a declared value becomes geometry, PACE measures, field by field, how far the compiled camera and the staged render sit from the declaration, rather than asking a model to judge.   On the 11-scene Automatic Drive screenplay, every staged single-subject panel places its subject within 1.2% of frame width of its declared position; with two or three subjects one camera pose cannot satisfy every position, and the residual is reported rather than absorbed. On 204 external director-storyboard shots, delivered head height is 1.906 times the staged target from the director&apos;s words, 1.733 from the compiled prompt, and 0.955 with the greybox control; the condition that holds framing best draws the described action least. Declaring the pose on 30 shots raises the action drawn from 58.9% to 74.4% without moving the framing. Transitions, fitted motion and human review of the generated panels remain open. Code: https://github.com/StudioPiLabs/pace-core</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19853v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19853v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Constraint-Safe Graph-Context Scoring for Stable Point-Feature Labels Under Text-Width and Accessibility-Inspired Profiles</title>
      <link>https://arxiv.org/abs/2609.19848v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19848v1</guid>
      <pubDate>Thu, 17 Sep 2026 08:00:00 GMT</pubDate>
      <dc:creator>Taimoor Ahmad</dc:creator>
      <category>模型架构</category>
      <description>Point-feature label placement on interactive maps must reconcile geometric validity, display yield, local placement utility, and stability across camera motion. Accessibility and multilingual requirem...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Constraint-Safe Graph-Context Scoring for Stable Point-Feature Labels Under Text-Width and Accessibility-Inspired Profiles</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Taimoor Ahmad</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Point-feature label placement on interactive maps must reconcile geometric validity, display yield, local placement utility, and stability across camera motion. Accessibility and multilingual requirements further change label dimensions, yet algorithmic evaluations often collapse these concerns into overlap counts. We present LABELSENSE-Pilot, a reproducible prototype that generates eight compass candidates per feature, scores candidates with a multilayer perceptron over graph-context summaries, adds a previous-placement bonus, and selects a layout through mixed-integer optimization. The executed scorer is deliberately not described as a graph transformer. Every returned layout is checked for viewport containment, per-feature uniqueness, and pairwise clearance. Experiments use 2,500 airport coordinates and names spanning 155 countries, with country-grouped splits and generated density, camera, text-suffix, preference, and enlarged-font stressors. Across five seeds, LABELSENSE-Pilot displayed 85.62 percent of labels with 2.09 percent flicker and zero collisions. Versus a handcrafted-utility integer program, LABELSENSE-Pilot sacrificed 1.43 percentage points of display while reducing flicker by 12.04 points. Enlarged-box-aware layouts produced zero proxy violations, whereas standard geometry reevaluated at 1.5x violated 52.57 percent of selected placements. These results establish an auditable engineering trade-off, not human accessibility, multilingual usability, or preference. Official recent baselines and participant evidence remain required before submission.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19848v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19848v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies</title>
      <link>https://arxiv.org/abs/2609.19844v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19844v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:54:54 GMT</pubDate>
      <dc:creator>Hang Xiao, Chuhong Xu, Kainan Zhou et al.</dc:creator>
      <category>模型架构</category>
      <description>AI-generated RTL verification plans can satisfy a provider schema yet fail at the boundary to trusted execution. We present SecTB-RTL, an auditable framework covering 31 tasks and 124 authored hardwar...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Hang Xiao, Chuhong Xu, Kainan Zhou et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">AI-generated RTL verification plans can satisfy a provider schema yet fail at the boundary to trusted execution. We present SecTB-RTL, an auditable framework covering 31 tasks and 124 authored hardware-security regressions. A deterministic non-AI baseline killed 36, 75, and 78 mutants at increasing resource limits. The first confirmatory run (C1-R2) failed before model execution because the provider rejected its response schema. After a schema-only repair made without viewing outcomes, a separately frozen follow-up run (C1-R3) completed 1,860 calls. The provider accepted 1,857 responses, but only nine passed the production semantic validator. The generation and execution rules did not match. We therefore preserve the run as an instrument-validation incident and report no prompt-effect estimate. This incident shows that provider or schema acceptance does not establish execution validity. Compilation and coverage are only diagnostics; the exact saved artifact must pass the full production path. A subsequent follow-up is excluded because it did not satisfy the preregistered evidence-completeness gate and is treated only as future work. We release the benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being misreported as model behavior.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19844v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19844v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents</title>
      <link>https://arxiv.org/abs/2609.19843v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19843v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:54:15 GMT</pubDate>
      <dc:creator>Haya Halimeh, Sascha Kaltenpoth, Kevin Bösch et al.</dc:creator>
      <category>图像生成</category>
      <description>LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliber...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Haya Halimeh, Sascha Kaltenpoth, Kevin Bösch et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in the textual outputs of LLMs are well-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions---and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it. Drawing on Dual-Process Theory, we empirically investigate whether LLM-based GUI agents are susceptible to automatic (Type 1) and reflective (Type 2) digital nudges, and how their reasoning configuration moderates this susceptibility. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect. Exploratory analysis further showed this redirection to be systematically structured by model scale. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19843v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19843v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks</title>
      <link>https://arxiv.org/abs/2609.19842v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19842v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:52:27 GMT</pubDate>
      <dc:creator>Shiyue Su, Song Wang, Zekai Zhan et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Effective EEG decoding requires representations that preserve organization among channels, local waveform dynamics, and long-range temporal context. Existing EEG architectures often capture these stru...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shiyue Su, Song Wang, Zekai Zhan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Effective EEG decoding requires representations that preserve organization among channels, local waveform dynamics, and long-range temporal context. Existing EEG architectures often capture these structures using separate specialized modules or collapse them into a single token sequence, making it difficult to maintain their distinct roles and coordinate their interactions throughout the backbone. We propose TriDim, a reusable block that preserves the representation shape and keeps three EEG axes explicit: channel, sample position within each patch, and patch position across the recording. These axes correspond to spatial, short-term temporal, and long-term temporal information, respectively. Each TriDim block applies feed-forward transformations along individual axes and cross-axis attention to coordinate information exchange among them. By stacking TriDim blocks with a multi-level tri-axis readout, we construct TriDimEEG, a standalone EEG decoder. Under strict cross-subject evaluation on eight datasets spanning clinical diagnosis, sleep staging, motor imagery, and emotion recognition, TriDimEEG achieves the best overall performance among fifteen evaluated models, with a 4.3% relative improvement in average accuracy over the second-best model. Replacing Transformer blocks in three EEG foundation models with TriDim blocks yields an average relative improvement of 7.4% in downstream accuracy while reducing parameter counts by 17.0% to 47.3%. These results establish TriDim as an effective and reusable building block and TriDimEEG as a strong standalone EEG decoder. Code and parameters of TriDimEEG are available at https://github.com/ncclab-sustech/TriDim_model.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19842v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19842v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark</title>
      <link>https://arxiv.org/abs/2609.19840v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19840v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:46:37 GMT</pubDate>
      <dc:creator>Hyunjung Chung, Unsang Park</dc:creator>
      <category>模型架构</category>
      <description>High-quality 3D talking face datasets remain largely English- centric, and Korean 3D facial motion data are difficult to combine with standard English benchmarks because of differences in mesh topolog...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Hyunjung Chung, Unsang Park</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">High-quality 3D talking face datasets remain largely English- centric, and Korean 3D facial motion data are difficult to combine with standard English benchmarks because of differences in mesh topology, spatial scale, coordinate system, and temporal sampling. We present KoUniTalk, a lightweight articulation-centered Korean-English 3D talk- ing face benchmark that retargets VOCASET and the released Korean speech-based 3D talking face data to a shared mesh topology using de- formation transfer. Rather than proposing a new deformation-transfer algorithm or a full-head identity-preserving avatar dataset, KoUniTalk provides an identity-neutral canonical output space for controlled speech- driven facial articulation training and evaluation across English and Ko- rean. The unified template contains 1,176 vertices and focuses on the mouth and adjacent lower- and mid-face regions, reducing the output dimensionality from 15,069 and 72,147 dimensions to 3,528 dimensions, corresponding to 4.27-fold and 20.45-fold reductions compared with VO- CASET/FLAME and the original Korean mesh, respectively. To exam- ine whether retargeting preserves speech-relevant motion, we evaluate semantic mouth-landmark trajectories, including mouth opening, mouth width, aperture ratio, and mouth-opening dynamics. Since the official test set of the Korean dataset is not publicly released, we additionally define a subject-disjoint Korean benchmark split. The processed matched benchmark contains 22 speakers, 4,978 sequences, and 642,781 frames, enabling Korean-English cross-dataset evaluation of speech-driven 3D fa- cial animation models in a single compact articulation-template space. Source-reported inventory counts are listed separately from these pro- cessed counts</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19840v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19840v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles</title>
      <link>https://arxiv.org/abs/2609.19831v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19831v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:36:34 GMT</pubDate>
      <dc:creator>Noah Mamié, Laurin van den Bergh</dc:creator>
      <category>模型架构</category>
      <description>In this reproducibility study, we investigate the transparency and scrutability of recommender systems enhanced by incorporating generated natural-language user profiles that represent user preference...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Noah Mamié, Laurin van den Bergh</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">In this reproducibility study, we investigate the transparency and scrutability of recommender systems enhanced by incorporating generated natural-language user profiles that represent user preferences. The original paper explores the synthesis of user profiles from raw user-generated review text across domains such as movies and accommodations (Amazon Movies &amp; TV, TripAdvisor). Crucially, these natural-language user profiles enable direct user interaction and intervention, allowing users to customize recommendations by correcting misattributed preferences or addressing cold-start settings. We successfully reproduce the core findings of the original study. Additionally, we extend the evaluation by conducting systematic context ablation experiments, multi-seed stability across five distinct random seeds to establish statistical reliability, and a mechanistic interpretability analysis using the nnsight framework to probe internal model representations under counterfactual profile perturbations. Our findings verify the original paper&apos;s claim that User Profile Recommendation (UPR) achieves competitive performance under its test-set reranking protocol and makes recommendations more transparent. Perturbing the natural-language profiles does change predictions, but it shifts predicted ratings uniformly across genres with no detectable genre-selective effect, leaving rankings unchanged even under direct activation steering. We trace this back to the rating-regression objective rather than the profile interface, with ranking-objective models clearly exceeding in this task.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19831v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19831v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization</title>
      <link>https://arxiv.org/abs/2609.19830v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19830v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:35:02 GMT</pubDate>
      <dc:creator>Yingxuan Zhuang, Binhe Yu, Jingxiao Yang et al.</dc:creator>
      <category>模型架构</category>
      <description>Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yingxuan Zhuang, Binhe Yu, Jingxiao Yang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19830v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19830v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] PyStream: Enhancing Video Streaming Evaluation</title>
      <link>https://arxiv.org/abs/2609.19823v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19823v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:31:37 GMT</pubDate>
      <dc:creator>Samuel Radler, Leon Prüller, Emanuele Artioli et al.</dc:creator>
      <category>模型架构</category>
      <description>As streaming services become more commonplace, analyzing their behavior effectively under different network conditions is crucial. This is normally quite expensive, requiring multiple players with dif...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PyStream: Enhancing Video Streaming Evaluation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Samuel Radler, Leon Prüller, Emanuele Artioli et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">As streaming services become more commonplace, analyzing their behavior effectively under different network conditions is crucial. This is normally quite expensive, requiring multiple players with different bandwidth configurations to be emulated by a powerful local machine or a cloud environment. Furthermore, emulating a realistic network behavior or guaranteeing adherence to a real network trace is challenging. This paper presents PyStream, a simple yet powerful way to emulate a video streaming network, allowing multiple simultaneous tests to run locally. By leveraging a network of Docker containers, many of the implementation challenges are abstracted away, keeping the resulting system easily manageable and upgradeable. We demonstrate how PyStream not only reduces the requirements for testing a video streaming system but also improves the accuracy of the emulations with respect to the current state-of-the-art. On average, PyStream reduces the error between the original network trace and the bandwidth emulated by video players by a factor of 2-3 compared to Wondershaper, a common network traffic shaper in many video streaming evaluation environments. Moreover, PyStream decreases the cost of running experiments compared to existing cloud-based video streaming evaluation environments such as CAdViSE.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19823v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19823v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[评估与优化] Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy</title>
      <link>https://arxiv.org/abs/2609.19820v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19820v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:30:23 GMT</pubDate>
      <dc:creator>Luis Leal</dc:creator>
      <category>评估与优化</category>
      <description>Regularized self-play -- the family behind DeepNash&apos;s Stratego play -- drives a two-player zero-sum policy to a Nash equilibrium by best-responding to a slowly moving, entropy-regularized reference po...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Luis Leal</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 评估与优化</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> rlhf</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Regularized self-play -- the family behind DeepNash&apos;s Stratego play -- drives a two-player zero-sum policy to a Nash equilibrium by best-responding to a slowly moving, entropy-regularized reference policy $ρ$. When the game has a polytope of value-equivalent equilibria, the regularizer silently breaks the tie: with a uniform reference it selects the maximum-entropy member, the I-projection of $ρ$ onto the Nash set. Can the reference be used to choose the equilibrium on purpose? On five exactly solvable games plus a 2-D polytope, with exact best responses and equivalence tests over independent seeds, anchoring the reference at a target member and refining steers self-play to that member with mean coordinate error 0.007 at median exploitability $5\times10^{-5}$, TOST-equivalent to the request within $\pm0.05$; the anchoring persists through refinement and follows the reference, not the initialization. Selection follows the reach-weighted I-projection (slope 0.969 [0.950, 0.987]). We report with equal emphasis where the story breaks: fixed off-manifold references cost 0.08-0.25 exploitability; stiff or flat families require a smaller mirror step, set by a pre-registered rule; boundary targets undershoot; curvature predicts where boundary saturation bites (rank correlation 0.90, p=0.037) while interior precision is curvature-independent. Table and MLP steering maps are equivalent within $\pm0.03$ at every target (30 seeds); matched control arms show attention&apos;s robust signature is excess seed variance, any systematic shift bounded at 0.018 and not significant. Against a best response the selection-robustness trade-off is degenerate: steering matters only against fixed, non-equilibrium opponents. The recipe -- anchor the reference at the desired member and refine -- reinterprets the KL anchor of RLHF-style RL as a selection knob, not only a stability leash.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19820v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19820v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 评估与优化 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection</title>
      <link>https://arxiv.org/abs/2609.19818v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19818v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:28:40 GMT</pubDate>
      <dc:creator>Kunyu Feng, Yuxiang Wang, Li Wang et al.</dc:creator>
      <category>模型架构</category>
      <description>Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an alrea...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kunyu Feng, Yuxiang Wang, Li Wang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by adapting recurrent inputs to the frozen encoder, controlling state updates, and aligning refined outputs with the frozen classifier. By training only lightweight refinement modules and loop-specific low-rank adapters on the original data, CoReLoop enables additional refinement while preserving the detector&apos;s original first-pass prediction. On 14 cross-domain test sets, the 24-layer model reduces pooled equal error rate (EER) from 4.85% to 3.74% with two passes, with approximately 10M trainable parameters out of 598M. To selectively apply this refinement, an optional halting head chooses the depth for each utterance, achieving 3.73% pooled EER with an average of 1.18 passes.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19818v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19818v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes</title>
      <link>https://arxiv.org/abs/2609.19815v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19815v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:26:56 GMT</pubDate>
      <dc:creator>Suji Kang, Seok-Young Kim, Young Bin Kim et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>We propose SnapPhysics, a training-free framework that reconstructs 3D objects and estimates their physical properties such as mass, friction, and center of gravity from a single image. For physically...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Suji Kang, Seok-Young Kim, Young Bin Kim et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We propose SnapPhysics, a training-free framework that reconstructs 3D objects and estimates their physical properties such as mass, friction, and center of gravity from a single image. For physically coherent interactions in mixed reality (MR), such properties are as important as geometry. Prior approaches infer them by analyzing object dynamics in video, which is computationally costly, or by querying vision-language models (VLMs) on single images, which lacks geometric grounding and inter-object relationships. We address these limitations by combining instance-level 3D reconstruction and spatial alignment with a physics-aware scene graph that encodes these relationships and per-object metric geometry as structured context for VLM-based property reasoning. Experiments on 3D-FRONT show that SnapPhysics improves scene-level F-Score by 18.6% over the best learning-based method, and on real captured scenes with ground-truth mass, it reduces the mean absolute log difference error (mALDE) by up to 20.5% and improves log-scale correlation ($r^2_{\mathrm{ls}}$) by up to 19.6% over VLM-only estimation. SnapPhysics enables physically interactive MR experiences without manual parameter tuning. Project page: https://snapphysics-ismar2026.github.io/.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19815v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19815v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Long-horizon autoformalization of a core theorem underlying MIP* = RE</title>
      <link>https://arxiv.org/abs/2609.19814v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19814v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:26:49 GMT</pubDate>
      <dc:creator>Sirui Lu, Ruixuan Deng, Yanqiao Zhu et al.</dc:creator>
      <category>模型架构</category>
      <description>Landmark mathematical formalizations have taken specialist teams years to complete. We present FormalFlow, a system that coordinates AI proving agents under human supervision to address statement drif...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Long-horizon autoformalization of a core theorem underlying MIP* = RE</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sirui Lu, Ruixuan Deng, Yanqiao Zhu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Landmark mathematical formalizations have taken specialist teams years to complete. We present FormalFlow, a system that coordinates AI proving agents under human supervision to address statement drift and proof composition in long-horizon formalization. Drawing on software engineering principles and practices, it uses a shared blueprint to guide nested planning, proving and review loops. Agents strengthen verification and review throughout formalization. We completed a machine-checked Lean 4 proof of the quantum soundness of the classical low individual-degree test, a core theorem underlying MIP* = RE. Developing the proof took 63 days; greater parallelism could further reduce this time. The final library contains 126,367 lines of Lean code, all generated by agents. The formalization corrects side conditions and intermediate errors while preserving the published final error bound under corrected assumptions. This work provides a verified foundation for quantum complexity and demonstrates a route to affordable verification of major research proofs by small teams.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19814v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19814v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint</title>
      <link>https://arxiv.org/abs/2609.19812v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19812v1</guid>
      <pubDate>Thu, 17 Sep 2026 07:24:40 GMT</pubDate>
      <dc:creator>Zhiyun Jiang, Hanyong Wang, Binbin Liang et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually a...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhiyun Jiang, Hanyong Wang, Binbin Liang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, vision-language model, vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually absent is vital for risk mitigation. To bridge this gap, we focus on visual scene negative captioning with safety as the cognitive constraint. The core challenge is to convert physical absence into semantic negative events. Existing vision-language models (VLMs) struggle with this process because affirmation bias suppresses negative reasoning, while limited mental filling capability and representation bias further hinder the inference of absent information. To address these challenges, we propose a negative captioning framework based on counterfactual reconstruction and contrastive decoding (CRCD). Inspired by human cognition, CRCD reformulates the task as counterfactual latent change captioning to bypass affirmation bias. It contrasts a synthesized safe expectation with reality to identify semantic omissions. To address limited mental filling, we design a dual-branch counterfactual reconstruction architecture. The amodal completion branch restores defective objects, while the functional association branch infers completely absent safety objects. Concurrently, a multi-condition representation learning mechanism is integrated to mitigate representation bias by projecting universal features onto predefined safety criteria subspaces, thereby capturing information across more dimensions. By decoding feature-level semantic residuals between the reconstructed scene prototype and raw input, CRCD bounds the non-existence search space and activates the decoder&apos;s negative logic. Extensive experiments validate the effectiveness of CRCD, establishing a high-performance baseline for this pioneering task.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19812v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19812v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agentic Personalization</title>
      <link>https://arxiv.org/abs/2609.19793v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19793v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:59:58 GMT</pubDate>
      <dc:creator>Xu Yuan, Yi Wang, Zhuohang Jiang et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Recent advances in artificial intelligence (AI) are reshaping smart glasses from egocentric capture and display devices into platforms for wearable intelligence. Smart glasses increasingly serve as we...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agentic Personalization</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xu Yuan, Yi Wang, Zhuohang Jiang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, personalization, gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Recent advances in artificial intelligence (AI) are reshaping smart glasses from egocentric capture and display devices into platforms for wearable intelligence. Smart glasses increasingly serve as wearable AI systems that connect first-person observation with real-time assistance under strict form-factor constraints. We frame this transition through the lens of \emph{AI smart glasses} and define them as a system-level concept in which egocentric sensing, resource-aware computing, intelligent reasoning, multimodal interaction, and real-world application constraints are co-designed for personalized assistance in the physical world. To systematically study this perspective, we organize the survey around four connected dimensions. First, we examine the hardware foundation that bounds sensing, computation, feedback delivery, and sustained deployment. Second, we study wearable intelligence, where egocentric signals are transformed into perceptual, contextual, and agentic capabilities. Third, we discuss interaction design, through which users request, receive, correct, and regulate assistance during ongoing activity. Fourth, we analyze application scenarios across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support, showing how domain requirements reshape system design and evaluation. We further identify five cross-cutting research challenges for future AI smart glasses: next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models. By centering smart glasses as wearable-intelligence platforms, this survey provides a unified framework for organizing technologies, applications, and open challenges in this emerging area.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19793v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19793v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements</title>
      <link>https://arxiv.org/abs/2609.19786v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19786v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:52:51 GMT</pubDate>
      <dc:creator>Caterina Amendola, Giulia Maffeis, Lorenzo Buffoni et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>The inverse problem of reconstructing optical properties, specifically absorption and scattering coefficients, in layered biological media from time-domain reflectance measurements remains a significa...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Caterina Amendola, Giulia Maffeis, Lorenzo Buffoni et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The inverse problem of reconstructing optical properties, specifically absorption and scattering coefficients, in layered biological media from time-domain reflectance measurements remains a significant challenge for traditional analytical models. Inverse solvers based on the diffusion equation often struggle with structural heterogeneity, frequently yielding poor accuracy for superficial absorption and deep-layers scattering. In this work, we propose a machine learning framework as an alternative approach to reconstruct the optical properties of a bilayered medium, benchmarking its efficiency and accuracy against model-based algorithms. To overcome the intrinsic approximations of diffusion theory and inverse reconstruction, we generated a robust synthetic dataset of forward DTOF using exact Monte Carlo simulations at multiple source-detector distances. A machine learning pipeline was then trained on this dataset and validated against state-of-the-art model-based reconstruction methods. Besides the significant reconstruction speed-up, the machine learning approach achieves higher accuracy than model-based inverse solvers, further providing an estimate of the parameter space dimensionality without requiring any a priori information about the number of layers in the investigated geometry. Further enhancements in the reconstruction accuracy can be expected in future extensions of this work, by training the pipeline over multiple DTOF curves from the same medium, in a joint multi-distance reconstruction approach.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19786v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19786v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Printing the Underdetermined: Materializing Multi-solutionness in Figurative Paintings</title>
      <link>https://arxiv.org/abs/2609.19782v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19782v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:50:54 GMT</pubDate>
      <dc:creator>Yutao Ming, Teng Xu, Youjia Wang et al.</dc:creator>
      <category>模型架构</category>
      <description>Figurative paintings are often approached as if they depict a single recoverable 3D scene: viewers infer depth and occlusion, and reconstruction pipelines attempt to converge to one stable model. We i...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Printing the Underdetermined: Materializing Multi-solutionness in Figurative Paintings</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yutao Ming, Teng Xu, Youjia Wang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Figurative paintings are often approached as if they depict a single recoverable 3D scene: viewers infer depth and occlusion, and reconstruction pipelines attempt to converge to one stable model. We instead foreground multi-solutionness, the non-uniqueness of 3D configurations compatible with a single painted image, and propose a workflow that keeps this non-uniqueness visible and material. Multi-solutionness arises from two sources: unobserved content, where backsides and occluded volumes admit multiple plausible completions, and observed cues, where perspective, shading, and occlusion still underconstrain geometry. When additional views are synthesized by a video generative model without explicit 3D constraints, small frame-level drifts become inevitable rather than exceptional. Our pipeline samples multiple camera-orbit multi-view video sequences from one painting, reconstructs each sequence with 3D Gaussian Splatting into a point-based Gaussian scene representation where density halos and ghosting expose unresolved degrees of freedom, and fabricates these representations as physical artifacts using DreamPrinting. By treating multiple compatible interpretations as explicit outputs rather than residual error, we provide a computational framework for spatial readings of figurative painting that can be inspected, compared, and discussed in both digital and physical form.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19782v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19782v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] PhyRestore: Physics-Structured Latent-Factor Restoration</title>
      <link>https://arxiv.org/abs/2609.19776v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19776v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:43:18 GMT</pubDate>
      <dc:creator>Ahmed Shafee, Chayan Lahiri</dc:creator>
      <category>模型架构</category>
      <description>Estimating temporal soil-loss change is challenging when physically meaningful input factors are noisy or corrupted, particularly because substantial changes are rare relative to the large number of l...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PhyRestore: Physics-Structured Latent-Factor Restoration</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ahmed Shafee, Chayan Lahiri</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Estimating temporal soil-loss change is challenging when physically meaningful input factors are noisy or corrupted, particularly because substantial changes are rare relative to the large number of locations exhibiting little change. We study this problem through the Revised Universal Soil Loss Equation (RUSLE) and introduce PhyRestore, a physics-structured latent-factor restoration framework. Rather than directly predicting soil-loss change or correcting a degraded physical estimate, PhyRestore restores corrupted physical factors and reconstructs temporal change through the known physical relationship. We evaluate PhyRestore in a watershed-scale bitemporal raster setting under isolated and simultaneous corruption of rainfall erosivity and cover management, comparing it with the degraded RUSLE estimate and Direct RF, XGBoost, MLP, and CNN models. Factor restoration improves high-magnitude recovery when the corrupted factors remain identifiable, but its advantage weakens under joint corruption, sparse positive extremes, and factor values outside the training support.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19776v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19776v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Integrating knowledge from case reports: a medical ontology based multimodal information system with structured summary</title>
      <link>https://arxiv.org/abs/2609.19775v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19775v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:41:18 GMT</pubDate>
      <dc:creator>Shuyu Guo, Lan Huang, Yichen Liu et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Published medical case reports serve as a crucial medical information carrier, documenting discoveries in rare diseases, diagnostic methods, and innovative treatments. Despite the wealth of clinical k...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Integrating knowledge from case reports: a medical ontology based multimodal information system with structured summary</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shuyu Guo, Lan Huang, Yichen Liu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Published medical case reports serve as a crucial medical information carrier, documenting discoveries in rare diseases, diagnostic methods, and innovative treatments. Despite the wealth of clinical knowledge in millions of case reports in the public medicine literature database (PubMed), accessing relevant information efficiently is hindered by the limitations of traditional keyword-based retrieval tools on unstructured and diverse case reports. To address the above issues, we introduce a comprehensive multimodal information system for case reports integrating structured clinical summaries of patients including medical images and biomedical named entities from 52949 open-access case reports published from 2000 to 2021. The multimodal essential information is organized in a well-structured medical ontology. Also, a powerful interface for searching and browsing case reports is designed to assist junior clinicians in retrieving cases effectively and improving the identification and diagnosis of rare diseases.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19775v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19775v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] TorchCraft: Unified binder design by inverting an all-atom structure predictor</title>
      <link>https://arxiv.org/abs/2609.19770v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19770v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:37:20 GMT</pubDate>
      <dc:creator>TorchCraft Team, Yu Liu, Zhouhanyu Shen et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>All-atom structure predictors model diverse molecular interactions, but using their learned structural priors for binder design remains challenging. Here we present TorchCraft, a unified binder-design...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">TorchCraft: Unified binder design by inverting an all-atom structure predictor</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> TorchCraft Team, Yu Liu, Zhouhanyu Shen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">All-atom structure predictors model diverse molecular interactions, but using their learned structural priors for binder design remains challenging. Here we present TorchCraft, a unified binder-design framework that optimizes sequence logits through a frozen all-atom predictor. Implemented in TorchFold, TorchCraft combines confidence, contact, geometric, and sequence-prior objectives within a shared optimization procedure for minibinders, framework-conditioned VHHs, cyclic peptides, and ligand-binding proteins. Using pretrained AlphaFold 3 weights, TorchCraft generated representative minibinders and VHHs with experimentally measured binding across four targets in each format, without post hoc sequence redesign. Computational benchmarks further demonstrated the framework&apos;s applicability to cyclic peptides and ligand-conditioned pocket design. TorchCraft extends predictor inversion to multiple binder formats and molecular contexts, providing a common framework for reusing all-atom structural priors in design.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19770v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19770v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting</title>
      <link>https://arxiv.org/abs/2609.19768v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19768v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:36:39 GMT</pubDate>
      <dc:creator>Yishun Zhu, Jian Wang</dc:creator>
      <category>模型架构</category>
      <description>Multivariate ocean forecasting must exploit shared evolution in a coupled ocean system while adapting to the heterogeneous statistical and dynamical characteristics of different prediction variables a...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yishun Zhu, Jian Wang</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Multivariate ocean forecasting must exploit shared evolution in a coupled ocean system while adapting to the heterogeneous statistical and dynamical characteristics of different prediction variables and locations. Fully shared models may lack the flexibility to handle this heterogeneity, whereas fully independent models discard the common ocean context shared across variables. The key question is how to retain shared context in a unified model while allowing computation to specialize according to the prediction target and local state. We propose OceanMoE, a structured conditional sparse Mixture-of-Experts framework that combines sharing and specialization for multivariate ocean forecasting. OceanMoE fuses cross-variable information to construct target-specific local representations and uses them to perform content-conditioned sparse routing at each spatial location, with the number of active experts adapted to router confidence. In the decoder, routing is augmented with a learned geographic bias parameterized by spherical-harmonic spatial bases, while shared residual and seasonal pathways provide common cross-variable and month-dependent context. Experiments on long-horizon autoregressive ORAS5 forecasting show that OceanMoE lowers aggregate forecasting error in both evaluated settings and maintains lower geometric-mean normalized RMSE than the corresponding baselines over most later rollout months. Routing analyses further show that expert allocation varies with prediction targets and spatial locations. These results support structured conditional computation as a modeling strategy for balancing shared ocean context with adaptive specialization.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19768v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19768v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding</title>
      <link>https://arxiv.org/abs/2609.19767v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19767v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:36:32 GMT</pubDate>
      <dc:creator>Zhiyun Jiang, Hanyong Wang, Binbin Liang et al.</dc:creator>
      <category>模型架构</category>
      <description>True machine intelligence requires transcending passive pixel registration to master top-down functional reasoning over absent information via visual negation understanding. However, unconstrained vis...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhiyun Jiang, Hanyong Wang, Binbin Liang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">True machine intelligence requires transcending passive pixel registration to master top-down functional reasoning over absent information via visual negation understanding. However, unconstrained visual negation paradigms remain overly open-ended, and pervasive affirmation bias causes both existing Multi-Modal Large Language Models (MLLMs) and evaluation metrics to fail under negative semantics. To solve these intertwined challenges systematically, we first anchor the boundaries of negation reasoning within specific cognitive goals. Specifically, by focusing on safety as a highly pragmatic and critical cognitive dimension, we define the task of \textbf{S}cene \textbf{N}egation \textbf{U}nderstanding under \textbf{S}afety Cognition (\textbf{SNUS}). Under this framework, we construct a high-fidelity negative caption dataset mapping dense assertions of localized hazards. Concurrently, we propose the Cognitive Expected Scene Graph (CESG) Score, a structure-grounded, polarity-aware evaluation metric. Extensive experiments demonstrate that while current models struggle on the task, traditional metrics completely collapse under semantic reversals. Conversely, our framework delivers a solid benchmark for SNUS, providing a rigorous foundation to advance risk-aware situational comprehension and counterfactual cognition.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19767v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19767v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] HyperAMS-Net: Adaptive Multi-Scale Spatial Hypergraph Network for Brain Disorder Classification</title>
      <link>https://arxiv.org/abs/2609.19755v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19755v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:25:21 GMT</pubDate>
      <dc:creator>Proloy Kumar Mondal, Md Kamran Hussin Chowdhury, Hoi Leong Lee</dc:creator>
      <category>模型架构</category>
      <description>Accurate classification of brain disorders from neuroimaging data remains challenging because of substantial inter-subject heterogeneity and the complex multi-scale patterns present in functional conn...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">HyperAMS-Net: Adaptive Multi-Scale Spatial Hypergraph Network for Brain Disorder Classification</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Proloy Kumar Mondal, Md Kamran Hussin Chowdhury, Hoi Leong Lee</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Accurate classification of brain disorders from neuroimaging data remains challenging because of substantial inter-subject heterogeneity and the complex multi-scale patterns present in functional connectivity and morphological representations. To address these challenges, we propose HyperAMS-Net, a deep learning framework for brain disorder classification using neuroimaging representations derived from resting-state functional MRI or structural MRI. HyperAMS-Net integrates adaptive multi-scale convolution, hypergraph attention, spatial-channel attention, and adaptive feature fusion. Specifically, adaptive multi-scale convolution learns data-driven weights over multiple receptive fields to capture complementary patterns at different scales. Hypergraph attention models higher-order dependencies among learned feature representations through node--hyperedge--node message passing, while spatial-channel attention enhances discriminative feature learning. Adaptive feature fusion further aggregates complementary information across parallel network branches. HyperAMS-Net is evaluated on three benchmark datasets spanning distinct brain disorders: ABIDE for autism spectrum disorder, REST-meta-MDD for major depressive disorder, and ADNI for Alzheimer&apos;s disease, using 5-fold stratified cross-validation. HyperAMS-Net achieves state-of-the-art performance across all evaluated datasets, attaining the highest accuracy and AUC among the compared methods. Ablation studies further demonstrate the contribution of each proposed component, with the largest performance degradation observed when hypergraph attention is removed.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19755v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19755v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] AutoData: Agentic Search for Pre-training Data Selection</title>
      <link>https://arxiv.org/abs/2609.19754v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19754v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:24:52 GMT</pubDate>
      <dc:creator>Yan Meng, Dhruv Srikanth, Bingchen Zhao et al.</dc:creator>
      <category>模型架构</category>
      <description>LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optim...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AutoData: Agentic Search for Pre-training Data Selection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yan Meng, Dhruv Srikanth, Bingchen Zhao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19754v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19754v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] STAR: Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction</title>
      <link>https://arxiv.org/abs/2609.19747v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19747v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:18:08 GMT</pubDate>
      <dc:creator>Wontae Choi, Ki Ryum Moon, Jae Young Lee et al.</dc:creator>
      <category>扩散模型</category>
      <description>Light field (LF) reconstruction from limited and noisy focal stack (FS) measurements is a highly ill-posed inverse problem. Although the LF-to-FS imaging geometry is fixed for a given optical setup, L...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">STAR: Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Wontae Choi, Ki Ryum Moon, Jae Young Lee et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Light field (LF) reconstruction from limited and noisy focal stack (FS) measurements is a highly ill-posed inverse problem. Although the LF-to-FS imaging geometry is fixed for a given optical setup, LF spatial-angular structure---including within-view spatial details, cross-view angular dependencies, and disparity across views---varies across scenes. Consequently, a fixed pre-trained prior may not optimally capture the spatial-angular structure of each test LF. We propose Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction (STAR), the first test-time adaptation framework for reconstructing an LF from FS. For each test LF, STAR freezes a pre-trained diffusion prior and fits three lightweight adapters to the observed FS to jointly adapt the three components of the LF&apos;s spatial-angular structure. STAR outperforms existing state-of-the-art methods in both two- and three-focal-sheet settings, with shorter inference times than those with test-time parameter updates.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19747v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19747v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Region-Level Policy Optimization for Fine-grained MLLM Perception</title>
      <link>https://arxiv.org/abs/2609.19745v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19745v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:10:34 GMT</pubDate>
      <dc:creator>Yuheng Shi, Xiaohuan Pei, Minjing Dong et al.</dc:creator>
      <category>模型架构</category>
      <description>Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two op...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Region-Level Policy Optimization for Fine-grained MLLM Perception</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yuheng Shi, Xiaohuan Pei, Minjing Dong et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have different resolution requirements. In a controlled diagnostic, localization tolerates roughly 3 to 4 times stronger token compression than recognition, which motivates localizing from a coarse view and concentrating resolution on the selected evidence. Decoding coordinates with the MLLM can be trained end-to-end from answers, but costs a full model pass per query and depends on grounding ability. A lightweight proposal network distilled from the model&apos;s attention is fast, but inherits the noise of its attention targets. The RoI from the proposal network reaches the answer through a discrete region choice, so its faithfulness to the answer cannot supervise the network. We therefore optimize the proposal network with region-level reinforcement learning, which we call Vision-RL2. It treats coherent regions as actions, and a frozen MLLM reader scores each one by how its removal changes the answer likelihood. Complementary subtractive and additive objectives suppress distracting proposals and recover missing evidence, updating only the predictor without region annotations, response sampling, or reasoning trajectories. The refined proposal further enables a sparse encoding that magnifies evidence and excludes background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens. Code is available at https://github.com/YuHengsss/VisionRL2 .</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19745v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19745v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Federated Learning Framework for Privacy-Preserving Kidney Stone Detection</title>
      <link>https://arxiv.org/abs/2609.19740v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19740v1</guid>
      <pubDate>Thu, 17 Sep 2026 06:06:03 GMT</pubDate>
      <dc:creator>Najiyya Younas, Omar Abdulkader, Yaser Ali Shah et al.</dc:creator>
      <category>图像生成</category>
      <description>Recent innovations in deep learning have significantly enhanced the diagnosis of medical images, although they are based on the use of centralized data storage that pose severe threats to patient priv...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Federated Learning Framework for Privacy-Preserving Kidney Stone Detection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Najiyya Younas, Omar Abdulkader, Yaser Ali Shah et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Recent innovations in deep learning have significantly enhanced the diagnosis of medical images, although they are based on the use of centralized data storage that pose severe threats to patient privacy and medical data security. To address this issue, this research proposes a Federated Learning (FL) model that is coupled with an optimized YOLOv8 network to detect the kidney stones on a computed tomography (CT) image and at the same time, protect privacy of the patients. The suggested system can help various medical organizations to jointly train a common model without exchanging the information about the patients. This is to ensure that data protection laws like GDPR and HIPAA are adhered to. The residual feature fusion and DropBlock regularization among other architectural improvements are also included in YOLOv8 to enhance detection robustness and minimize overfitting. Experimental analysis carried out on a distributed CT dataset demonstrated that the federated YOLOv8 model has a mAP at 50 of 0.733 and is able to keep the data confidential. Moreover, its lean design facilitates fast edge deployment and real-time inference across a clinical setting. Altogether, these findings indicate that Federated Learning is a safe and efficient solution to AI-assisted diagnosis in contemporary healthcare when combined with the use of sophisticated object detection models.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19740v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19740v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] The segmentation ceiling: why explicit left-ventricular masks do not improve learned ejection-fraction regression</title>
      <link>https://arxiv.org/abs/2609.19730v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19730v1</guid>
      <pubDate>Thu, 17 Sep 2026 05:41:28 GMT</pubDate>
      <dc:creator>Farshid Farhadi Khouzani, Paul La Plante, Bryar Mustafa Shareef et al.</dc:creator>
      <category>模型架构</category>
      <description>Accurate estimation of left ventricular ejection fraction (EF) from echocardiography is central to cardiovascular care, and deep learning enables automated EF prediction from echocardiographic video. ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">The segmentation ceiling: why explicit left-ventricular masks do not improve learned ejection-fraction regression</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Farshid Farhadi Khouzani, Paul La Plante, Bryar Mustafa Shareef et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> mae</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Accurate estimation of left ventricular ejection fraction (EF) from echocardiography is central to cardiovascular care, and deep learning enables automated EF prediction from echocardiographic video. Because EF is clinically derived from left-ventricular (LV) volumes, a widely held intuition is that explicit LV segmentation should improve prediction. We introduce a quantitative criterion, the segmentation ceiling, that makes this testable: from EF as a normalized difference of end-diastolic and end-systolic volumes, we derive in closed form how per-frame segmentation area error propagates into EF error, and thus the accuracy a mask must reach before it can improve on direct regression. Using EchoNet-Dynamic, a UniFormer-S backbone, and the empirically measured within-patient error correlation, the criterion places the break-even near 10% per-frame area error, whereas a representative segmenter operates at roughly 14%, above the ceiling. Consistent with this, four strategies for injecting segmentation or area information (a predicted-mask channel, end-diastolic/end-systolic clip sampling, and per-bin and amplitude area-consistency objectives) fail to beat a raw-video baseline; ground-truth masks help only through label leakage. Input representation thus not being the limit, we identify generalization as the practical lever: weight averaging with strong augmentation attains a test R^2 of 0.806 (MAE 4.08) under a matched dense-clip protocol, comparable to an R(2+1)D baseline (0.811) while tightening the validation-to-test gap. Finally, a heteroscedastic beta-NLL formulation yields informative, well-calibrated per-prediction uncertainty, larger for clinically harder low-EF cases, where Monte-Carlo dropout does not. The segmentation ceiling gives a concrete design criterion for when mask-guided EF estimation is worthwhile, plus a simple, uncertainty-aware recipe for EF regression.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19730v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19730v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [模型架构] Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation</title>
      <link>https://arxiv.org/abs/2609.19729v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19729v1</guid>
      <pubDate>Thu, 17 Sep 2026 05:40:53 GMT</pubDate>
      <dc:creator>Tri Cao, Hung Nguyen, Phong Nguyen et al.</dc:creator>
      <category>视频生成</category>
      <category>模型架构</category>
      <description>Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames res...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tri Cao, Hung Nguyen, Phong Nguyen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, video generation, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response $R( Δ, \, t_{\text{denoise}})$, a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from $R$, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19729v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19729v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents</title>
      <link>https://arxiv.org/abs/2609.19721v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19721v1</guid>
      <pubDate>Thu, 17 Sep 2026 05:27:46 GMT</pubDate>
      <dc:creator>Meysam Ghaffari, Bhaskar Sen, Nasim Sabetpour et al.</dc:creator>
      <category>模型架构</category>
      <description>Clinical coding agents repeatedly encounter the same failure modes, including unsupported codes, missed documented conditions, specificity errors, and procedure-coding convention mismatches. We introd...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Meysam Ghaffari, Bhaskar Sen, Nasim Sabetpour et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Clinical coding agents repeatedly encounter the same failure modes, including unsupported codes, missed documented conditions, specificity errors, and procedure-coding convention mismatches. We introduce Learn-Then-Act, an inference-time adaptation framework that converts errors from a small labeled LEARN batch into a structured Mistake Knowledge Database (MistakeKDB). False-negative lessons are routed to a recall-oriented Coder, while false-positive lessons are routed to a precision-oriented Judge. We instantiate the framework in LearnActCoder, a Coder-Judge clinical coding pipeline with lookup-table grounding where available. On 150 matched MIMIC-III notes, structured MistakeKDB improves CPT F1 by 5.9 percentage points, while raw-example and reflection-style memories remain near the no-memory baseline; the ICD-9 improvement is not significant. On a matched MIMIC-IV cohort, memory shifts ICD-10 coding toward higher precision at a recall cost, leaving F1 statistically unchanged. Applying the same memory to 1,000 held-out MIMIC-III notes maintains a stable ICD operating point, providing scale/stability evidence. Overall, the results are consistent with structured, feedback-derived error memory being useful for adapting clinical coding behavior across cases without weight updates or changes to the underlying workflow. Absolute CPT/HCPCS performance remains low, and the system is evaluated retrospectively rather than in clinical deployment.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19721v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19721v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] GAPrompt++: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model</title>
      <link>https://arxiv.org/abs/2609.19716v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19716v1</guid>
      <pubDate>Thu, 17 Sep 2026 05:14:34 GMT</pubDate>
      <dc:creator>Zixiang Ai, Zhenyu Cui, Yufei Guo et al.</dc:creator>
      <category>模型架构</category>
      <description>Pre-trained 3D vision models have substantially advanced point cloud analysis, yet adapting them to downstream tasks via full fine-tuning is computationally expensive and storage-intensive. Parameter-...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">GAPrompt++: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zixiang Ai, Zhenyu Cui, Yufei Guo et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Pre-trained 3D vision models have substantially advanced point cloud analysis, yet adapting them to downstream tasks via full fine-tuning is computationally expensive and storage-intensive. Parameter-Efficient Fine-Tuning (PEFT) offers a promising alternative by reducing both adaptation cost and storage burden. However, existing prompting-based approaches ignore the intrinsic geometric structures of point clouds, thereby limiting their adaptation capability. This limitation stems from their inability to encode both fine-grained geometric cues and coarse-grained structural semantics, as well as failing to propagate such information effectively through the model hierarchy. To address these challenges, we propose GAPrompt++, a multi-granular geometry-aware prompting method that provides richer geometric guidance for efficient 3D task adaptation. Specifically, we introduce a Point Shift Prompter that extracts multi-granular geometric features across different scales, enabling instance-specific geometric adjustments during adaptation. Next, a Keypoint Prompter adaptively generates point-level prompts to highlight local geometric saliency and fine-grained structural details. Furthermore, a Prompt Propagation mechanism injects these multi-granular geometric cues throughout the feature extraction hierarchy, strengthening the ability to capture essential geometric characteristics. Extensive experiments show that GAPrompt++ achieves state-of-the-art performance among prompting-based PEFT methods and even surpasses full fine-tuning across diverse benchmarks, while requiring less than 2\% trainable parameters. In addition, to address the saturation of existing evaluation datasets, we construct two more challenging benchmarks derived from 3D Gaussian Splatting and Multi-View Stereo reconstruction, offering diverse and realistic point cloud scenarios to promote future research.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19716v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19716v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits</title>
      <link>https://arxiv.org/abs/2609.19709v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19709v1</guid>
      <pubDate>Thu, 17 Sep 2026 05:02:37 GMT</pubDate>
      <dc:creator>Sulgi Kim</dc:creator>
      <category>模型架构</category>
      <description>Batched multi-armed bandits update on a service&apos;s own schedule, and the usual implementation carries each arm&apos;s absolute reward rate from one update to the next. When the shared level moves between ba...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sulgi Kim</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Batched multi-armed bandits update on a service&apos;s own schedule, and the usual implementation carries each arm&apos;s absolute reward rate from one update to the next. When the shared level moves between batches, that memory goes stale even though the comparisons between arms may not have. Odds-Ratio Thompson Sampling (OR-TS) instead carries the joint posterior over log-odds contrasts and fits the common level afresh in every batch, marginalizing it out. This paper specifies that update, places it inside a Bayesian bandit agent with two controls, decay for how much past evidence survives an update and aggressiveness for how sharply belief becomes allocation, and evaluates it against absolute-rate memory. Across 86 public A/B series the level varies about twenty-five times more than the contrast. In prespecified synthetic environments a moving level costs absolute-rate memory five times the regret and leaves the best arm below a majority of traffic in 7 of 20 runs, against none for OR-TS. In a policy simulation built from 71 real experiments, where the contrasts are too small to resolve, expected-click differences stay within 0.1% for 58 of them, yet contrast memory still ends on the better arm more than twice as often. Where the contrasts themselves move, the bet fails, and that case is reported too.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19709v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19709v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation</title>
      <link>https://arxiv.org/abs/2609.19702v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19702v1</guid>
      <pubDate>Thu, 17 Sep 2026 04:54:59 GMT</pubDate>
      <dc:creator>Daeun Kim, Junwha Hong, Changhun Oh et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. However, generating thousands of visual...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Daeun Kim, Junwha Hong, Changhun Oh et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> image generation, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. However, generating thousands of visual tokens per request makes decoding increasingly bottlenecked by KV cache accesses during attention computation. Sparse attention is particularly attractive for this workload because many visual generation applications tolerate moderate quality degradation in exchange for improved performance and efficiency. While sparse attention has been extensively explored for text-based LLM inference, it remains unclear whether its sparsity assumptions generalize effectively to autoregressive image generation. We present the first systematic characterization of attention sparsity in autoregressive image generation across diverse workloads and representative open-source models. Our analysis reveals several distinguishing properties, including a pronounced prefill-decode asymmetry, strong attention concentration on prompt and local tokens, and a unique diagonal attention sparsity pattern arising from the spatial locality of visual tokens. Motivated by these observations, we propose a diagonal-aware sparse attention mechanism that selectively skips KV entries along the diagonal attention direction within a recent window. Implemented on top of a GPU-based serving system using FlexGen, FlashAttention-2, and custom kernels, our approach achieves up to 3.1x throughput and 1.19x latency improvements with less than 2% quality degradation compared to dense inference.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19702v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19702v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems</title>
      <link>https://arxiv.org/abs/2609.19695v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19695v1</guid>
      <pubDate>Thu, 17 Sep 2026 04:44:06 GMT</pubDate>
      <dc:creator>Mohammed El Hanjri, Anas Abouaomar, Hamidou Tembine et al.</dc:creator>
      <category>模型架构</category>
      <description>Federated learning (FL) enables privacy-preserving, on-device training across heterogeneous Internet-of-Things (IoT) deployments such as smart-city water-metering networks, where each smart meter obse...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Mohammed El Hanjri, Anas Abouaomar, Hamidou Tembine et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> mae, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Federated learning (FL) enables privacy-preserving, on-device training across heterogeneous Internet-of-Things (IoT) deployments such as smart-city water-metering networks, where each smart meter observes a household-specific consumption time series. Under such statistical heterogeneity, the standard Federated Averaging (FedAvg) aggregation averages dissimilar local models into a single global model that may fail to capture client-specific patterns. We address this by forming client coalitions directly in the local-weight space and aggregating at the coalition level. Extending a prior weight-driven coalition-formation scheme, we model coalition formation as a Hegselmann-Krause (HK) bounded-confidence opinion-dynamics process acting on the local weights, and develop variants of the HK interaction based on Euclidean-distance and cosine-similarity confidence criteria. The framework is applied to short-term water-consumption forecasting with local Long Short-Term Memory (LSTM) models and evaluated against FedAvg, Per-FedAvg, FedProx, and FedAvg with Euclidean-distance or cosine-similarity coalition formation. Experiments on a real smart-metering dataset of water consumption show that the proposed HK-based coalition formation produces stable, endogenous coalition structures within at most ten inner iterations, incurs no additional client-side computation or communication compared to FedAvg, and reduces the average MAE by up to 54% relative to FedAvg, 39% relative to FedProx, and 24% relative to Per-FedAvg, while achieving the highest global accuracy (83-85%).</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19695v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19695v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models</title>
      <link>https://arxiv.org/abs/2609.19693v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19693v1</guid>
      <pubDate>Thu, 17 Sep 2026 04:40:20 GMT</pubDate>
      <dc:creator>Dasom Choi, Sangjun Moon, Hyeongchan Im et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring backgrou...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Dasom Choi, Sangjun Moon, Hyeongchan Im et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring background context and inter-face relationships, which often yields suboptimal performance. To overcome these limitations, we leverage instruction-based Large Vision-Language Models (LVLMs), which can interpret entire images and follow complex textual instructions. We propose a simple yet effective single-stage multi-face forgery detector, called IMFD (Instruction-based Multi-face Forgery Detector), which is trained end-to-end to jointly localize faces and predict per-face forgery labels. Rather than treating face box prediction only as a joint objective, IMFD explicitly integrates predicted face bounding boxes into the instruction as visual cues that enhance instruction grounding and forgery detection. To support the training and evaluation of IMFD, we convert existing multi-face forgery datasets into an instruction-based format. Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19693v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19693v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] UniExo: Unified Multi-Skill Policies for Musculoskeletal Locomotion and Co-Adaptive Exoskeleton Control</title>
      <link>https://arxiv.org/abs/2609.19690v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19690v1</guid>
      <pubDate>Thu, 17 Sep 2026 04:38:34 GMT</pubDate>
      <dc:creator>Yifei Yuan, Jakob Wolf, Ghaith Androwis et al.</dc:creator>
      <category>模型架构</category>
      <description>Daily locomotion encompasses diverse activities and frequent transitions between them, yet most exoskeleton controllers are designed for a single activity or a narrow set of related movements. Changes...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">UniExo: Unified Multi-Skill Policies for Musculoskeletal Locomotion and Co-Adaptive Exoskeleton Control</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yifei Yuan, Jakob Wolf, Ghaith Androwis et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Daily locomotion encompasses diverse activities and frequent transitions between them, yet most exoskeleton controllers are designed for a single activity or a narrow set of related movements. Changes in activity therefore typically require explicit mode switching and separately tuned or retrained controllers. Simulation-based learning reduces the need for hardware-based tuning but generally retains this limitation. Here we present UniExo, a framework that first constructs a multi-skill musculoskeletal human policy and then jointly trains an exoskeleton control policy with it. Four single-skill imitation experts for walking, turning, running and backward walking are distilled into a single network structured by a skill latent and subsequently fine-tuned through reinforcement learning on transition sequences. The resultant unified human policy achieves a mean tracking success rate of 94.7% on unseen clips of the four skills and exhibits greater robustness to perturbations than its constituent experts. A single hip exoskeleton controller (UniExo) is initialized from hip moment prediction of the human policy and co-adapted with it through multi-agent reinforcement learning across the four skills. This co-adaptation shifts the timing of the assistance torque and raises the fraction of positive work delivered to the hip. When deployed on a custom hip exoskeleton, the controller generalizes across four treadmill speeds in six participants and assists one participant through a continuous route of all four skills and their transitions, without skill labels or explicit mode switching. UniExo thus provides a step towards replacing activity-specific controllers with unified, user-specific controllers that support diverse locomotor activities and the transitions between them.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19690v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19690v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration</title>
      <link>https://arxiv.org/abs/2609.19683v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19683v1</guid>
      <pubDate>Thu, 17 Sep 2026 04:31:34 GMT</pubDate>
      <dc:creator>Yuan Liao, Jae-sun Seo</dc:creator>
      <category>多模态生成</category>
      <description>The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerators are strictly cons...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yuan Liao, Jae-sun Seo</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerators are strictly constrained by area and power, they require end-to-end quantized models. However, the extreme dynamic range gap between multi-modal tokens causes standard block formats to suffer &quot;microscaling collapse,&quot; where a single massive outlier hijacks the shared exponent, underflowing surrounding elements and destroying attention maps. To break this bottleneck, we propose Micro-Inverted-Scaling (MiX), a novel format that mathematically inverts the microscaling paradigm: rather than grouping multiple mantissas under one shared exponent, MiX groups private, per-element exponents under a single shared mantissa. To handle asymmetric VLM outlier topologies, we introduce an adaptive dual-format (MiX-MX) inference framework. By algebraically factoring out the shared MiX mantissa, this framework maps to a custom accelerator, replacing multipliers with efficient shifters. Evaluated end-to-end on multiple VLMs, our 4.5-bit MiX formulation exhibits equivalent or superior accuracy on multi-modal benchmarks compared to NVFP4. Simultaneously, the MiX accelerator delivers a 25% improvement in area efficiency over the NVFP4 baseline and a 2.3-4.5x speedup with 1.4-2.9x energy reduction across models compared to the state-of-the-art accelerator Focus, proving the inverted-scaling datapath is physically superior for efficient VLM deployment.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19683v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19683v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models</title>
      <link>https://arxiv.org/abs/2609.19674v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19674v1</guid>
      <pubDate>Thu, 17 Sep 2026 04:16:51 GMT</pubDate>
      <dc:creator>Yufeng Wang, Parivesh Priye, Lu Wei et al.</dc:creator>
      <category>模型架构</category>
      <description>A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yufeng Wang, Parivesh Priye, Lu Wei et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follow the law seen during training rather than the intervened one. We show that these two failures require different structural remedies. Evolving a learned energy with a symplectic integrator preserves the geometry of the conservative dynamics and keeps rollouts bounded and physically meaningful for up to $100\times$ the training horizon, while equal-capacity predictors, an energy-regularized predictor, and a tuned neural ODE diverge. By contrast, encoding the physical coupling through an explicit linear factorization enables the model to follow a never-seen sign of that coupling, whereas an unrestricted parameterization remains locked to the training law. Crucially, the two mechanisms are separable: removing the structure responsible for long-horizon stability leaves counterfactual transfer intact, while removing the factorized coupling destroys counterfactual transfer without eliminating stability. This double dissociation, established with matched controls that remove or replace one structural component at a time, persists beyond the headline three-body system and remains visible when the physical state must be inferred from pixels rather than provided directly. The result is a concrete design principle for physical world models: long-horizon stability and changed-law generalization arise from distinct structural commitments, and each can be imposed deliberately without requiring the other.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19674v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19674v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[评估与优化] When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models</title>
      <link>https://arxiv.org/abs/2609.19671v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19671v1</guid>
      <pubDate>Thu, 17 Sep 2026 04:15:27 GMT</pubDate>
      <dc:creator>Jaejun Shim, HyunJin Kim, Young Jin Kim et al.</dc:creator>
      <category>评估与优化</category>
      <description>Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jaejun Shim, HyunJin Kim, Young Jin Kim et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 评估与优化</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> reward model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on easy instances for accuracy loss on hard instances. We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism that leverages pre-computed reference statistics (accuracy and token usage) to regulate reasoning depth. Combined with verifier-based rewards and batch-wise standardized advantages, IDAC enables stable critic-free optimization without learned reward models or online reference-model queries. When2Think encourages direct answering on easy instances while preserving extended reasoning on hard instances, thereby learning when to use System 1 (NoThink) versus System 2 (Think). Experiments on mathematical benchmarks demonstrate improved accuracy-efficiency trade-offs: on AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model, and on AIME25, When2Think achieves 40.0% Pass@3, outperforming compression and routing-only baselines.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19671v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19671v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 评估与优化 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] [图像生成] EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence</title>
      <link>https://arxiv.org/abs/2609.19659v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19659v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:59:11 GMT</pubDate>
      <dc:creator>Feifan Wang, Zongbing Zhang, Yu Zhang et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilizati...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Feifan Wang, Zongbing Zhang, Yu Zhang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, edm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training paradigm that achieves state-of-the-art average performance through strategic data selection and hierarchical policy optimization. Our approach consists of three synergistic stages. First, Rejection Sampling-based Fine-Tuning (RSFT) filters out low-informative samples to establish robust behavioral priors while preventing distributional collapse. Second, Iterative Rejection GRPO (IR-GRPO) employs task-specific queues stratified by difficulty to keep datasets balanced across reinforcement learning iterations, coupled with a hybrid reward mechanism for precise cross-task feedback. Third, to enhance long-horizon task planning, we introduce Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees, which enables step-level advantage estimation. This resolves the credit assignment problem by isolating intermediate correct decisions from downstream errors, while effectively balancing exploration efficiency and depth compared to conventional search trees. As a result, EmbodiedMind achieves a state-of-the-art average performance of 70.02% across 18 benchmarks, and significantly outperforms other embodied foundation models in long-horizon task planning accuracy. Our project will be released for reproducibility.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19659v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19659v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving</title>
      <link>https://arxiv.org/abs/2609.19657v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19657v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:56:20 GMT</pubDate>
      <dc:creator>Omkar Shewale, Deepak Kumar, Divakar Kumar Yadav</dc:creator>
      <category>模型架构</category>
      <description>Repeated prompt prefixes are increasingly common in LLM serving workloads, appearing in system prompts, templated retrieval-augmented generation pipelines, agent frameworks, and multi-turn conversatio...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Omkar Shewale, Deepak Kumar, Divakar Kumar Yadav</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Repeated prompt prefixes are increasingly common in LLM serving workloads, appearing in system prompts, templated retrieval-augmented generation pipelines, agent frameworks, and multi-turn conversations. Modern inference runtimes such as vLLM and TensorRT-LLM provide mechanisms for reusing previously computed KV-cache state across requests, yet it remains unclear when prefix reuse materially improves serving performance on contemporary accelerators and when its benefits are limited by scheduling, cache granularity, concurrency, or memory pressure.   This paper presents PrefixBench-H100, a reproducible benchmark and measurement framework for characterizing prefix reuse on a single NVIDIA H100. PrefixBench-H100 combines controlled synthetic traces with chat-style and retrieval-style workloads, and evaluates two widely used LLM serving runtimes under matched workload conditions. The benchmark varies shared-prefix length, suffix diversity, request arrival pattern, concurrency, output length, and cache configuration, while collecting time-to-first-token, inter-token latency, end-to-end latency, throughput, cache-hit statistics, GPU memory usage, and selected profiling traces.   The goal of PrefixBench-H100 is not to introduce a new caching algorithm, but to expose the practical operating envelope of prefix reuse for H100-class LLM serving. The study identifies the regime where prefix reuse provides substantial first-token latency reductions and the regime where cache pressure erodes them, while showing that cache effectiveness itself is largely insensitive to concurrency and output length; the cross-runtime differences that remain arise above the cache, in the scheduling layer.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19657v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19657v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Self-Evolving Search Index</title>
      <link>https://arxiv.org/abs/2609.19656v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19656v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:56:03 GMT</pubDate>
      <dc:creator>Sangam Lee, Wonjae Lee, Sunghwan Kim et al.</dc:creator>
      <category>模型架构</category>
      <description>Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Self-Evolving Search Index</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sangam Lee, Wonjae Lee, Sunghwan Kim et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19656v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19656v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions</title>
      <link>https://arxiv.org/abs/2609.19654v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19654v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:50:39 GMT</pubDate>
      <dc:creator>Xiaofei Yuan, Yan Zhang, Shaobo Qiao et al.</dc:creator>
      <category>模型架构</category>
      <description>Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involv...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xiaofei Yuan, Yan Zhang, Shaobo Qiao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison. We conduct a systematic empirical study using two TREK-derived benchmark sets: 500 single-disruption cases, including feasible and infeasible instances, and 200 feasible simultaneous compound-disruption cases. We compare LLM-Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO local-revision adapter across effectiveness, plan stability, and computational cost. LLM-Z3 with Gemini achieved the highest observed compound-disruption success. IPyHOPPER nearly matched that configuration&apos;s single-disruption overall success, while preserving substantially more of the accepted itinerary on successful repairs. Successful hierarchical and local repairs made fewer edits and retained more accepted commitments than full replanning. Computational profiles differed: IPyHOPPER used no LLM inference, the evaluated LLM-Z3 adapter used compact one-call inference, and the iTIMO adapter consumed substantially more tokens. The study provides practical guidelines for balancing feasibility recovery, commitment preservation, and computational cost within evaluated settings.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19654v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19654v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Well-posedness of neural turbulence closures and tangent dissipation</title>
      <link>https://arxiv.org/abs/2609.19647v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19647v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:40:41 GMT</pubDate>
      <dc:creator>Zhen Zhang, George Em Karniadakis</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>A neural turbulence closure defines a new boundary-value problem, $R(U)=N(U)+F(U)=0$, with a coupled Jacobian $J(U)=N&apos;(U)+F&apos;(U)$, where $N$ is the original mean-flow operator and $F$ the learned closu...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Well-posedness of neural turbulence closures and tangent dissipation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhen Zhang, George Em Karniadakis</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, diffusion</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">A neural turbulence closure defines a new boundary-value problem, $R(U)=N(U)+F(U)=0$, with a coupled Jacobian $J(U)=N&apos;(U)+F&apos;(U)$, where $N$ is the original mean-flow operator and $F$ the learned closure. We establish two consequences of global tangent dissipation. For a monotone original operator, a positive uniform margin supplied by the original operator and closure together guarantees existence, uniqueness and a global inverse-sensitivity bound relating a posteriori solution error to the a priori residual. For a general original operator, a dissipative closure cannot worsen tangent dissipation, but this alone does not guarantee uniqueness. Tangent dissipation depends on both diffusion and reaction. We study two complementary ways to promote it: (1) an exact-integral construction enforcing non-negative tangent diffusion while leaving reaction unconstrained, and (2) a penalty on tangent-reaction violations at sampled states. Tangent diffusion enters the Jacobian, and non-negative secant eddy viscosity alone does not control its coercivity. We conduct tests with channel flow at $Re_τ=180$--$5200$, which provides a strongly monotone baseline. Both constrained closures reach accurate solutions in all 50 training-seed/Reynolds-number cases. At $Re_τ=1000$, we conduct tests with 10,000 starts for one fixed network per closure and we find one root for each constrained closure and multiple roots for the other closures. Although this does not prove uniqueness, it provides strong empirical evidence for uniqueness of the tested constrained closures. At $Re_τ=5200$, the construction and penalty reduce the reported inverse sensitivity relative to the original operator by approximately $372\times$ and $11\times$, respectively.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19647v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19647v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A Policy Profile for Croissant: Refusal as a Property of the Dataset</title>
      <link>https://arxiv.org/abs/2609.19640v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19640v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:30:32 GMT</pubDate>
      <dc:creator>Alexander Chernov</dc:creator>
      <category>模型架构</category>
      <description>Croissant is the de facto machine-readable descriptor for ML datasets: JSON-LD over schema.org. Since version 1.1 it also carries data use conditions, recommending DUO and ODRL for them. What no versi...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Policy Profile for Croissant: Refusal as a Property of the Dataset</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Alexander Chernov</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Croissant is the de facto machine-readable descriptor for ML datasets: JSON-LD over schema.org. Since version 1.1 it also carries data use conditions, recommending DUO and ODRL for them. What no version specifies is how any of them is evaluated: no decision procedure, no bound on evaluation cost, no outcome for a condition an implementation cannot evaluate, no record of what was checked, and nothing on composition with caller-side authority. We supply that half. An additive profile lets a dataset declare the operations it admits and the conditions under which it admits them, over a closed set of five operators whose decision procedure is given in full, so a gate decides from the descriptor alone and records what it checked. Two corpora evaluate it and their evidence is kept apart. Three descriptors that gated a real nf-core pipeline give the deployment result: decisions from a profile document match the gate&apos;s native descriptor record for record, stripping the layer leaves a valid Croissant document, and the added cost is 11.7 $μ$s against a 119 $μ$s decision. A corpus generated from the profile&apos;s grammar gives the breadth, covering every operator, refusal class and conformance clause. Across its valid cases, 552 complete decision records agree three ways -- native descriptor, profile terms, and the same policy as ODRL in usageInfo. The carrier is therefore not the contribution; the evaluation semantics is. Finally, caller-bound and data-bound policies range over non-overlapping state spaces, so neither permit set contains the other.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19640v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19640v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs</title>
      <link>https://arxiv.org/abs/2609.19636v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19636v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:27:53 GMT</pubDate>
      <dc:creator>Xuan Liu, Jingbin Qian</dc:creator>
      <category>模型架构</category>
      <description>Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making. An agent in a closed loop ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xuan Liu, Jingbin Qian</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making. An agent in a closed loop writes its own inputs. Each observation follows from its own earlier actions, so the states it meets late in an episode are partly of its own making. An SFT checkpoint and an RL checkpoint are then scored from different states, even on identical tasks. Endpoint success mixes two changes: where the agent arrives, and what it does once it is there. Restricting the comparison to states both policies reach does not separate them. That restriction selects on an outcome, and in our data it flips the sign of the effect. We introduce checkpoint handoff, an evaluation protocol that clones a state one released checkpoint reached and hands it to another, with no retraining. Crossing a reacher role and a solver role over SFT and RL splits an endpoint gain into REACH and SOLVE. REACH is how often a policy arrives at a state the environment confirms is a fixed number of actions from success. SOLVE is how often it finishes from an identical cloned state. Across two benchmarks and two independently released pipelines, the reacher by solver interaction is positive in all five conditions. An RL history is worth more to an RL solver than the same history is to an SFT solver. On ALFWorld, RL improves both terms, and the SFT solver never succeeds where the RL solver fails. Independent REACH and SOLVE gaps predict the aggregate interaction. Handoff asks only that one checkpoint&apos;s history can be replayed under another, so long-horizon evaluation can report arrival and completion beside endpoint success.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19636v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19636v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Instance Segmentation and Fine-grained Classification for Urban Buildings with Adaptive Region Dividing and Spatially-Supervised Contrastive Learning</title>
      <link>https://arxiv.org/abs/2609.19631v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19631v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:23:54 GMT</pubDate>
      <dc:creator>Weiyuan Zhang, Qi Zhang, Hui Huang</dc:creator>
      <category>模型架构</category>
      <description>Accurate instance-level and functional understanding of urban buildings in large-scale point clouds is essential for digital city modeling and urban analysis. However, the extensive spatial coverage o...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Instance Segmentation and Fine-grained Classification for Urban Buildings with Adaptive Region Dividing and Spatially-Supervised Contrastive Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Weiyuan Zhang, Qi Zhang, Hui Huang</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Accurate instance-level and functional understanding of urban buildings in large-scale point clouds is essential for digital city modeling and urban analysis. However, the extensive spatial coverage of urban scenes leads most existing methods to rely on predefined blocks for training and evaluation, although such partitions are rarely available in real-world applications and introduce additional preprocessing while fragmenting complete building structures. To address this issue, we propose an adaptive region-dividing strategy with unified scene-level evaluation. Specifically, the 3D point cloud is projected onto a bird&apos;s-eye-view (BEV) plane, where a pretrained segmentation model is used to detect building regions. The detected bounding boxes are then back-projected to the original point cloud to construct structure-aligned adaptive training blocks, enabling semantically guided dynamic partitioning without manual design. Furthermore, beyond instance-level understanding, few methods have explored fine-grained classification for urban buildings, and thus we also put forward a fine-grained classification model for urban buildings with a spatially-supervised contrastive loss. First, for each segmented building, a point transformer classifier jointly encodes its body and local context using geometric, color, and core-context information. Then, the class-balanced weighted cross-entropy is used to alleviate severe class imbalance. The proposed spatially-supervised contrastive loss further enhances inter-class discriminability by assigning greater weight to spatially proximate, same-category buildings, encouraging compact functional representations while separating easily confused categories. Extensive experiments on UrbanBIS and STPLS3D demonstrate the advantages of the proposed method in building instance segmentation and fine-grained classification compared to existing SOTA methods.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19631v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19631v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] DataCanvas-EDU: An Agentic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education</title>
      <link>https://arxiv.org/abs/2609.19617v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19617v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:08:26 GMT</pubDate>
      <dc:creator>Bang An, Maria Hamdani, Joseph Fox</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Business analytics education requires diverse datasets to support different learning objectives, student backgrounds, and analytical tasks. Real-world data can be difficult to obtain and offer limited...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">DataCanvas-EDU: An Agentic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Bang An, Maria Hamdani, Joseph Fox</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Business analytics education requires diverse datasets to support different learning objectives, student backgrounds, and analytical tasks. Real-world data can be difficult to obtain and offer limited flexibility for adapting a case to a particular course. Even when suitable data are available, instructors must investigate the patterns, verify the results, and prepare assignments and reference solutions, requiring substantial time and effort. The use of large language models (LLMs) introduces an additional concern about training data contamination. Widely used public datasets often have extensive tutorials and worked analyses that models may have encountered during training. Students may therefore receive explanations drawn from existing analyses without practicing how to investigate unfamiliar data in collaboration with AI. This paper presents DataCanvas-EDU, an agentic framework for instructor-guided synthetic data generation in business analytics education. Instructors specify teaching goals and intended patterns through conversation, while an AI agent writes generation code, checks the resulting data, and prepares assignments, reference analyses, and rubrics. Four phases, Plan, Create, Verify / Test Analysis, and Evaluate, organize the process and support instructor review and revision. The framework is intended to simplify case preparation while creating opportunities for students to investigate newly designed patterns with AI. We illustrate the approach with WindowDash, a food delivery case containing 15,000 orders and nine designed patterns. DataCanvas-EDU is packaged as a reusable AI Agent Skill for compatible agent environments, with the package and installation instructions available at https://github.com/BANG23333/datacanvas-edu</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19617v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19617v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability</title>
      <link>https://arxiv.org/abs/2609.19616v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19616v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:07:39 GMT</pubDate>
      <dc:creator>Michael Hernandez, Tian Zhao</dc:creator>
      <category>模型架构</category>
      <description>Complexity measured from generated code is failure-dependent: a difficult prompt can yield a short failing program and be assigned low output complexity. We introduce a six-dimension prompt-side struc...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Michael Hernandez, Tian Zhao</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Complexity measured from generated code is failure-dependent: a difficult prompt can yield a short failing program and be assigned low output complexity. We introduce a six-dimension prompt-side structural-complexity index scored before generation and kept separate from correctness. We select 5,000 Python prompts across six bands of a preliminary single-rater rubric. Four out-of-panel LLM raters rescore the locked prompts, giving 19,997 score rows; composite inter-rater reliability is ICC = 0.872 on the 4,998 prompts with all four ratings. We evaluate 21 models per prompt, yielding 105,000 generations. In the unadjusted mean-pooled analysis, pass rate has a nonmonotone breakpoint at composite 13.75, with 79.9% at or below and 87.6% above. This is not a universal failure cutoff. Task-type fixed effects shift the breakpoint to 10.75 and cut the regime gap from 7.6 to 2.1 points. A construction-frame control shifts it to 8.50 with a raw gap of -3.5 points, and neither frame alone reproduces the pooled +7.6-point change. Model-specific fits include 16 upward and five downward changes. A 365-prompt audit-clean extension matches the original five-model estimates at bins 15 and 16 but adds only 14 prompts above bin 16. Among zero-pass generations with computable Lizard complexity, 28.5% pair a prompt composite above 8 with output complexity at most 10. Human agreement is moderate and rater-dependent on a disagreement-enriched calibration set; paraphrase and cross-language rescoring preserve score ordering. Overidentification tests reject the joint restrictions on the six dimensions, so we treat the composite as an index and make no causal interpretation of the 2SLS estimates. The contribution is a pre-generation measurement framework and a bounded observational analysis of reliability regimes.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19616v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19616v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation</title>
      <link>https://arxiv.org/abs/2609.19613v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19613v1</guid>
      <pubDate>Thu, 17 Sep 2026 03:01:35 GMT</pubDate>
      <dc:creator>Haodi Hu, Kaen Kogashi, Toshiaki Koike-Akino</dc:creator>
      <category>模型架构</category>
      <description>Dexterous food manipulation requires control under deformation, occlusion, and uncertain contact. We present TacSushi, a tactile-grounded, Cosmos3-based world-action policy that learns from recorded f...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Haodi Hu, Kaen Kogashi, Toshiaki Koike-Akino</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Dexterous food manipulation requires control under deformation, occlusion, and uncertain contact. We present TacSushi, a tactile-grounded, Cosmos3-based world-action policy that learns from recorded future consequences while acting on current observations. The backbone encodes current RGB, language, and hand state, and feature-wise gated fusion incorporates fingertip tactile features into the action representation. During training, a decoder conditioned on demonstrated action chunks predicts logged future visual observations, task progress, relative contact risk, and tactile summaries; this decoder is removed at deployment. Failed trials provide consequence supervision, but their actions are excluded from imitation. We train TacSushi on 340 successful and 50 failed real-robot trials and compare six methods in 600 separate rollouts across three in-distribution tasks and two out-of-distribution ingredient variants. To assess food quality beyond a single geometric threshold, we score terminal outcomes using an anchored visual-quality protocol that equally weights five human ratings and three vision-language-model ratings per rollout. Full TacSushi achieves 68.3% average in-distribution success and 37.5% out-of-distribution success, compared with 36.7%/10.0% without future-consequence supervision and 25.0%/17.5% with direct tactile concatenation in place of gated fusion. These comparisons support complementary benefits of feature-wise gated tactile fusion and training-only predictive supervision.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19613v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19613v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] The Output-Space Hypothesis: Enumerative Equivalence Checking for Tensor Programs</title>
      <link>https://arxiv.org/abs/2609.19611v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19611v1</guid>
      <pubDate>Thu, 17 Sep 2026 02:54:30 GMT</pubDate>
      <dc:creator>Paul Biberstein, Joseph Devietti, Mayur Naik</dc:creator>
      <category>模型架构</category>
      <description>Tensor programs, as used in deep learning models, are a prime target for optimization, as small performance improvements can have a large impact across training or inference workloads. However, such o...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">The Output-Space Hypothesis: Enumerative Equivalence Checking for Tensor Programs</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Paul Biberstein, Joseph Devietti, Mayur Naik</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Tensor programs, as used in deep learning models, are a prime target for optimization, as small performance improvements can have a large impact across training or inference workloads. However, such optimizations are complicated and can produce subtle bugs. Traditionally, correctness is assumed when differential testing against a reference on random inputs fails to reveal bugs. However, the inputs to these programs are massive tensors, and finding bugs can require generating extremely low likelihood inputs with precise relationships among their values.   We propose a novel way to find bugs more consistently by flipping the quantifiers. Rather than generating a single input and checking all output tensor locations for equivalence, what if you could check a single output tensor location&apos;s equivalence for all inputs? We implement this idea in a system, \dirigo, by using a novel symbolic execution strategy. We demonstrate that \dirigo can find bugs effectively in a public dataset of 6,988 AI-written CUDA kernels that are all marked correct by differential testing. Of these, \dirigo finds 600 kernels that are actually buggy, and finds 97.3\% of those bugs within two minutes.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19611v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19611v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership</title>
      <link>https://arxiv.org/abs/2609.19610v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19610v1</guid>
      <pubDate>Thu, 17 Sep 2026 02:43:54 GMT</pubDate>
      <dc:creator>Run Peng, Zinnia Nie, Jing Ding et al.</dc:creator>
      <category>图像生成</category>
      <description>Understanding humans over long horizons requires agents to infer not only what people need in the moment, but also how routines form, why they repeat, and when they change. We introduce SimLife, a sca...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Run Peng, Zinnia Nie, Jing Ding et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> personalization</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Understanding humans over long horizons requires agents to infer not only what people need in the moment, but also how routines form, why they repeat, and when they change. We introduce SimLife, a scalable platform for simulating long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues with audio. Built on SimLife, SimLife-BP evaluates long-context pattern understanding: the ability to infer latent behavioral rules from weeks or months of everyday observations. The benchmark contains 106 episodes averaging 15.49 hours and 38.57 in-game days, and 1,439 question-answer pairs. Each task probes direct, counterfactual, noisy, and inverse reasoning under different levels of rule hints. Evaluating frontier models and architectures, we find that current models often achieve surface-level prediction without comprehensive rule understanding, rely on frequency-based heuristics rather than if-then reasoning over evidence, and struggle to adapt when behavioral patterns change. These findings suggest that long-context pattern understanding remains a major bottleneck for future embodied agents, while SimLife opens a broader space for studying memory, personalization, adaptation, and long-horizon planning in everyday human-AI interaction.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19610v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19610v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction</title>
      <link>https://arxiv.org/abs/2609.19606v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19606v1</guid>
      <pubDate>Thu, 17 Sep 2026 02:40:31 GMT</pubDate>
      <dc:creator>Theodore O. Cochran</dc:creator>
      <category>图像生成</category>
      <description>This empirical study is an independent reproduction of the dissociation Zhao reported in 2026. The shape of a large language model&apos;s chain-of-thought entropy trajectory predicts whether the final answ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Theodore O. Cochran</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">This empirical study is an independent reproduction of the dissociation Zhao reported in 2026. The shape of a large language model&apos;s chain-of-thought entropy trajectory predicts whether the final answer is correct, while the magnitude of its total entropy drop does not. The dissociation merits reproduction because the magnitude half rests on a single 300-problem run with one model at one seed, while the shape half was reported at full scale on both benchmarks and on a second model family. Registered at OSF before any confirmatory run, the reproduction crosses the complete GSM8K and MATH-500 benchmark test sets with four open-weight models including one reasoning-distilled model of a kind the original did not test. The shape signal replicates. The magnitude signal divides by setting. On the anchor model the accuracy gap between monotone and non-monotone chains is +9.6 percentage points on GSM8K and +27.5 on MATH-500, while the rank correlation of the total entropy drop with correctness is -0.018 on GSM8K and +0.414 on MATH-500. On the reasoning-distilled model the binary form of the shape signal fires on about one chain in a hundred, too few to estimate the registered contrast, while the graded violation count remains predictive there. In an exploratory comparison the final-step entropy alone outperforms the binary shape flag in all eight model-by-benchmark cells by ROC area, and in six or seven by the risk-coverage area the original reports, depending on an integration range the original does not state. The study contributes a reproduction of the shape signal at full test-set scale under seven documented protocol differences, a map of the settings where the magnitude signal holds and fails, and measurements of four protocol dependencies the original does not report.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19606v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19606v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions</title>
      <link>https://arxiv.org/abs/2609.19592v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19592v1</guid>
      <pubDate>Thu, 17 Sep 2026 02:18:02 GMT</pubDate>
      <dc:creator>Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu et al.</dc:creator>
      <category>模型架构</category>
      <description>This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras und...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model (SAM), SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model (RAM). Among the detection models, GELAN-s achieved the most favorable balance between mean average precision (mAP) and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an $R^2$ value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19592v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19592v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Large Language Model Agents for Evidence Based Genetic Disease Severity Classification</title>
      <link>https://arxiv.org/abs/2609.19569v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19569v1</guid>
      <pubDate>Thu, 17 Sep 2026 01:56:38 GMT</pubDate>
      <dc:creator>Tohid Ghasemnejad, Ahmadreza Argha, Mark Grosser et al.</dc:creator>
      <category>模型架构</category>
      <description>Disease severity classification for genetic conditions is subjective and labor-intensive, creating bottlenecks in genomic screening, where commercial panels vary widely in size and overlap. We develop...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Large Language Model Agents for Evidence Based Genetic Disease Severity Classification</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tohid Ghasemnejad, Ahmadreza Argha, Mark Grosser et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Disease severity classification for genetic conditions is subjective and labor-intensive, creating bottlenecks in genomic screening, where commercial panels vary widely in size and overlap. We developed an autonomous AI agent integrating Reasoning and Acting (ReAct) with Retrieval-Augmented Generation (RAG) to classify 10,211 Human Phenotype Ontology terms. It uses American College of Medical Genetics (ACMG)-endorsed severity guidelines and American College of Obstetricians and Gynecologists (ACOG) quality-of-life criteria to retrieve PubMed literature, generate interpretable reasoning chains, and independently verify claims. At the phenotype level, using expert-curated cohorts, the agent achieved 93.55% accuracy (MCC 0.9237) with 82.6% to 91.4% of claims supported by direct evidence or valid inferences. Gene-level severity was aggregated across 8,738 pairs, identifying 3,283 autosomal recessive pairs with severe or profound presentations. External validation showed 95.2% concordance with Mackenzie&apos;s Mission gene list. This system enables standardized panel design by providing reliable, automated classification supported by direct evidence.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19569v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19569v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] FedFIbOS: Fisher Importance based Optimal Submodelling for Heterogeneous Federated Learning</title>
      <link>https://arxiv.org/abs/2609.19559v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19559v1</guid>
      <pubDate>Thu, 17 Sep 2026 01:41:17 GMT</pubDate>
      <dc:creator>Yasmeen Afzal, Jeremiah D. Deng, Haibo Zhang</dc:creator>
      <category>模型架构</category>
      <description>Heterogeneous federated learning requires clients with diverse computational capacities to collaboratively train a global model, where each client trains a capacity-constrained submodel. Existing meth...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">FedFIbOS: Fisher Importance based Optimal Submodelling for Heterogeneous Federated Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yasmeen Afzal, Jeremiah D. Deng, Haibo Zhang</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Heterogeneous federated learning requires clients with diverse computational capacities to collaboratively train a global model, where each client trains a capacity-constrained submodel. Existing methods select submodel parameters using heuristic importance measures---most prominently parameter magnitude---without theoretical justification for why these measures support convergence. We identify a fundamental gap: existing parameter selection criteria lack theoretical grounding in the convergence framework, partial client participation introduces additional estimation effects in the Fisher scores. We propose \textbf{FedFIbOS}: Fisher Importance-based Optimal Submodelling for heterogeneous federated learning, using Fisher Information in a principled criterion derived from minimizing submodel masking error. %We formally establish when magnitude selection is equivalent to Fisher selection fail under non-IID heterogeneous federated learning. We theoretically formulate submodel selection through a Fisher-weighted quadratic masking surrogate and show that the raw Fisher top-$k$ rule implemented by FedFIbOS solves this surrogate under a Fisher-dominant ranking condition. The resulting method retains the convergence structure of the underlying masked federated optimization bound. Fisher scores are efficiently estimated from empirical diagonal Fisher information using squared gradients, enabling stable and adaptive parameter selection without additional optimization overhead. Experiments on CIFAR-10, CIFAR-100, and AGNews under pathological and Dirichlet non-IID settings show FedFIbOS achieves ${\approx}10\%$ higher accuracy than the state of the art, with improvements becoming more pronounced under stronger heterogeneity.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19559v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19559v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control</title>
      <link>https://arxiv.org/abs/2609.19554v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19554v1</guid>
      <pubDate>Thu, 17 Sep 2026 01:24:51 GMT</pubDate>
      <dc:creator>Zhongbo Zhang, Jiayi Jin, Yifan Wang et al.</dc:creator>
      <category>模型架构</category>
      <description>Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhongbo Zhang, Jiayi Jin, Yifan Wang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19554v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19554v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Continual Enterprise World Model Discovery in Dynamic Systems</title>
      <link>https://arxiv.org/abs/2609.19551v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19551v1</guid>
      <pubDate>Thu, 17 Sep 2026 01:23:43 GMT</pubDate>
      <dc:creator>Shambhavi Mishra, David Vazquez, Perouz Taslakian et al.</dc:creator>
      <category>图像生成</category>
      <description>In an enterprise system, updating one field can set another, create a record, or start an approval. These effects are produced by business rules that are not built into the platform but written by eac...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Continual Enterprise World Model Discovery in Dynamic Systems</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shambhavi Mishra, David Vazquez, Perouz Taslakian et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">In an enterprise system, updating one field can set another, create a record, or start an approval. These effects are produced by business rules that are not built into the platform but written by each organization and revised over time. An agent working in such a system cannot predict the result of its own actions without knowing these rules. We study continual enterprise world model discovery, where an agent starts without knowledge of these business rules and discovers them by interacting with records and observing the outcomes. From those observations it builds a world model, which it revises as the rules change. To evaluate this, we introduce EnterpriseWorldShift, built on a live ServiceNow environment with nine tables, 25 hidden rules and 600 evaluation actions. It presents four versions of the same enterprise world, with the tables and records held fixed while a rule is modified, then added, then removed, so that discovery, revision, extension and retirement are each tested in turn. Our Continual Discovery Agent (CDA) builds such a model and carries it from one world to the next. It predicts the effects of the hidden rules more accurately than looking them up for each question, the approach taken by prior work, by up to 8.98 IoU points, and it answers from its own model without querying the running system.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19551v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19551v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] PerSeM: Persistent Semantic Memory for Long-Horizon Open-Vocabulary UAV Mapping</title>
      <link>https://arxiv.org/abs/2609.19542v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19542v1</guid>
      <pubDate>Thu, 17 Sep 2026 01:07:57 GMT</pubDate>
      <dc:creator>Saurbh Singh Jamwal, Ganesh Ramakrishnan</dc:creator>
      <category>模型架构</category>
      <description>Open-vocabulary segmentation enables rich semantic perception for UAVs, but frame-wise predictions can remain temporally inconsistent across repeated observations and changing viewpoints. We present P...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PerSeM: Persistent Semantic Memory for Long-Horizon Open-Vocabulary UAV Mapping</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Saurbh Singh Jamwal, Ganesh Ramakrishnan</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Open-vocabulary segmentation enables rich semantic perception for UAVs, but frame-wise predictions can remain temporally inconsistent across repeated observations and changing viewpoints. We present PerSeM, a training-free persistent semantic memory framework for long-horizon open-vocabulary UAV mapping. PerSeM associates frame-wise semantic observations with persistent world-space voxels and constructs a majority-based semantic memory, which is conservatively refined through history-preserving spatial refinement, trust-aware replay, and context-guided verification. Experiments on the Forest and UAVScenes benchmarks show that persistent 3D memory provides substantial gains in semantic correctness and temporal stability over frame-wise predictions. Beyond this strong persistent-memory baseline, PerSeM provides consistent additional improvements, improving both semantic accuracy and temporal stability across all five evaluated UAVScenes sequences. Analysis using regions identified independently of the final PerSeM predictions further shows that these gains are concentrated in semantically difficult and temporally unstable regions, where majority-based memory is most likely to remain uncertain. These results demonstrate that persistent 3D aggregation provides a strong foundation for long-horizon semantic mapping, while conservative refinement of uncertain memory states can provide additional improvements without retraining or additional neural-network inference.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19542v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19542v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks</title>
      <link>https://arxiv.org/abs/2609.19538v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19538v1</guid>
      <pubDate>Thu, 17 Sep 2026 01:04:27 GMT</pubDate>
      <dc:creator>Nguyen Duc Minh Quang, Chang Liu, Shuangyang Li et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Low-altitude wireless networks (LAWNs) are emerging as a key infrastructure for heterogeneous unmanned aerial systems that support concurrent services within a shared three-dimensional airspace. Their...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Nguyen Duc Minh Quang, Chang Liu, Shuangyang Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Low-altitude wireless networks (LAWNs) are emerging as a key infrastructure for heterogeneous unmanned aerial systems that support concurrent services within a shared three-dimensional airspace. Their coexistence creates strong coupling among mobility, connectivity, and shared network resources, while heterogeneous services impose distinct and time-varying requirements. These interactions naturally form a dynamic non-cooperative game in which both operating conditions and coordination objectives evolve over time. Conventional optimization and learning-based controllers typically rely on predefined objectives, limiting their ability to adapt autonomously to changing service requirements and resource priorities. To address this challenge, we propose a hierarchical hybrid large language model (LLM)- multi-agent reinforcement learning (MARL) architecture organized as a dual-loop structure. Specifically, an outer adaptation loop employs LLM-assisted game orchestration to interpret service requirements and operator intent, and reconfigure objectives and resource priorities, while an inner loop executes decentralized, parameter-conditioned MARL policies under the configured game. A logistics-monitoring case study illustrates how the proposed framework facilitates coordinated coexistence among heterogeneous services, adapting to evolving operating conditions without retraining the underlying MARL policies. Finally, we discuss key challenges and research directions toward scalable, trustworthy, and adaptive agentic LAWNs.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19538v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19538v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation</title>
      <link>https://arxiv.org/abs/2609.19527v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19527v1</guid>
      <pubDate>Thu, 17 Sep 2026 00:41:46 GMT</pubDate>
      <dc:creator>Keshu Wu, Hao Zhang, Rui Gan et al.</dc:creator>
      <category>模型架构</category>
      <description>Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execut...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Keshu Wu, Hao Zhang, Rui Gan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execute successfully while failing to realize the spatial, temporal, communication, or behavioral relationships requested by the user. This paper presents AURORA, a natural-language-driven agentic framework that treats air-ground scenario generation as a process of compilation with verification. Central to AURORA is the Air-Ground Scenario Graph (AGSG), a typed intermediate representation that explicitly connects agents, aerial missions, events, communication links, success conditions, and their cross-domain dependencies. This shared representation enables simulator-grounded parsing, joint road-airspace grounding, temporal planning, pre-execution feasibility checking, trace-based runtime verification, failure localization, and bounded repair within a unified workflow. We further introduce AURORA-Bench to evaluate not only whether generated scenarios execute, but whether they faithfully realize the requested interactions. Experiments across multiple language models show that structured execution substantially improves reliability, while runtime verification exposes silent failures that completion-based evaluation overlooks. Localized repair further resolves many violations without regenerating the entire scenario. The results show that reliable scenario generation requires verifying realized behavior, not merely executable code, and demonstrate the value of explicit intermediate representations for verifiable and repairable language-driven co-simulation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19527v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19527v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Self Improvement via Fast Tree-search</title>
      <link>https://arxiv.org/abs/2609.19526v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19526v1</guid>
      <pubDate>Thu, 17 Sep 2026 00:41:13 GMT</pubDate>
      <dc:creator>Xinghong Fu, Aravinth Kulanthaivelu, Yutaro Yamada</dc:creator>
      <category>图像生成</category>
      <description>Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are cost...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Self Improvement via Fast Tree-search</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xinghong Fu, Aravinth Kulanthaivelu, Yutaro Yamada</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-improvement framework that significantly improves coding performance under strict budget constraints. We identify evaluation of candidate self-modifications as the main runtime bottleneck since prior approaches estimate their effectiveness by re-running a subset of benchmark tasks with the modified agent, which is time-consuming. We introduce Recursive Self Improvement via Fast Tree-search (SIFT), which augments these downstream task evaluations with an LLM-as-a-judge signal that performs pairwise comparisons between candidate patches, where the win-loss record is aggregated with a regularized Bradley-Terry model, and the resulting strength scores drive rank-based parent sampling inside a lightweight disaggregated tree search. Expensive downstream task evaluations are reserved only for the most promising nodes. Using a fully disaggregated tree search pipeline, the judge scores provide intermediate signal to guide exploration on promising candidate patches without being bottlenecked by slow evaluation runs. SIFT outperforms existing tree-search based self-evolution frameworks on the full Polyglot benchmark with significantly lower resource requirements in terms of CPU hours, wall clock time, and API cost.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19526v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19526v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems</title>
      <link>https://arxiv.org/abs/2609.19524v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19524v1</guid>
      <pubDate>Thu, 17 Sep 2026 00:31:16 GMT</pubDate>
      <dc:creator>Shaina Raza, Ahmed Y. Radwan, Imran Liaquat et al.</dc:creator>
      <category>模型架构</category>
      <description>Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (ML...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shaina Raza, Ahmed Y. Radwan, Imran Liaquat et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessment, yet their evaluation evidence must remain interpretable for development and oversight. We propose a unified framework that connects output-level, trajectory-level, and cross-modal assessment through eight trustworthiness dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. The framework preserves system-specific metrics while mapping native measurements to common performance bands, accompanied by uncertainty estimates and traceable evidence. A meta-evaluation layer examines the validity, reliability, and reproducibility of the evaluation itself. Multidimensional profiles expose strengths and weaknesses, while safety-critical overrides prevent aggregate scores from masking critical failures. Mappings to governance frameworks, international standards, and European Union regulatory requirements connect technical assessment with oversight needs. The framework provides a structured basis for assessing both system performance and the credibility of the evidence supporting it, with empirical validation across deployment contexts remaining an essential next step.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19524v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19524v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] LSTM-UT and Recurrent-Depth Transformers on Cellular Automata</title>
      <link>https://arxiv.org/abs/2609.19521v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19521v1</guid>
      <pubDate>Thu, 17 Sep 2026 00:17:39 GMT</pubDate>
      <dc:creator>Aras Kavuncu</dc:creator>
      <category>模型架构</category>
      <description>Recurrent-depth Transformers apply shared computation repeatedly, but differ in how they retain information across steps. We compare a Block Universal Transformer (BUT), which carries only its current...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">LSTM-UT and Recurrent-Depth Transformers on Cellular Automata</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Aras Kavuncu</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Recurrent-depth Transformers apply shared computation repeatedly, but differ in how they retain information across steps. We compare a Block Universal Transformer (BUT), which carries only its current hidden state; CoTFormer, which also retains an expanding attention cache; and a new LSTM Universal Transformer (LSTM-UT) with bounded gated memory. On Rule 30 cellular automata, BUT extrapolates to unseen recurrent depths more reliably than CoTFormer, although its accuracy eventually degrades. State and cache interventions show that CoTFormer&apos;s failure depends on their interaction: correcting the current state can temporarily restore accuracy, while retained history can undermine that correction. In a delayed-recall task, BUT also outperforms CoTFormer despite lacking direct access to past states; CoTFormer does not reliably select the requested cached representation. LSTM-UT improves both depth extrapolation and delayed recall over these baselines. The results support bounded gated memory as an effective inductive bias for repeated computation and later retrieval in these tasks.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19521v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19521v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend</title>
      <link>https://arxiv.org/abs/2609.19518v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19518v1</guid>
      <pubDate>Thu, 17 Sep 2026 00:11:16 GMT</pubDate>
      <dc:creator>Hengyi Wang, Lourdes Agapito</dc:creator>
      <category>模型架构</category>
      <description>We present AMB3R-SLAM, a real-time monocular SLAM system capable of reconstructing kilometer-scale trajectories over 10k frames on a single consumer-grade GPU. Our model couples a lightweight front-en...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Hengyi Wang, Lourdes Agapito</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We present AMB3R-SLAM, a real-time monocular SLAM system capable of reconstructing kilometer-scale trajectories over 10k frames on a single consumer-grade GPU. Our model couples a lightweight front-end for low-latency online tracking with a hierarchical backend that progressively enforces local, mid-level, and global consistency. By avoiding bundle adjustment that relies on the static world assumption, our system naturally handles complex dynamic scenes out of the box. Furthermore, we demonstrate that our method can be extended to leverage stereo, RGB-D, and LiDAR as additional inputs. AMB3R-SLAM achieves strong camera tracking performance across 9 datasets, reducing the absolute trajectory error (ATE) of previous state-of-the-art methods on VBR and Oxford Spires by over 70%. With additional LiDAR input, our model further reduces ATE to sub-meter level on KITTI and VBR datasets.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19518v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19518v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] LLM-as-an-Improver: Turning Verification into Better Candidates</title>
      <link>https://arxiv.org/abs/2609.19515v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19515v1</guid>
      <pubDate>Thu, 17 Sep 2026 00:05:25 GMT</pubDate>
      <dc:creator>Akiyoshi Tomihari, Yuma Ichikawa</dc:creator>
      <category>模型架构</category>
      <description>Verifier-based selection improves LLM performance by generating multiple candidate solutions and using a verifier to select the most promising one. However, existing methods typically treat verificati...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">LLM-as-an-Improver: Turning Verification into Better Candidates</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Akiyoshi Tomihari, Yuma Ichikawa</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Verifier-based selection improves LLM performance by generating multiple candidate solutions and using a verifier to select the most promising one. However, existing methods typically treat verification only as a ranking step and discard its feedback once a fixed candidate pool has been evaluated. In this paper, we ask whether verification can also improve the candidate set itself. To this end, we introduce LLM-as-an-Improver and propose Verify--Repair--Reselect (VRR), which uses verification feedback to generate and reselect improved candidates. VRR retains the initial winner while conditionally generating three complementary alternatives: repaired versions of the winner and runner-up, and a solution based on a new approach. It filters invalid and duplicate candidates using only inference-time information and then reselects the final answer under the original evaluation criteria. Across diverse models and code-generation and reasoning benchmarks, VRR improves over fixed-pool verifier-based selection in many settings and can recover correct solutions even when all candidates in the initial pool are incorrect. These results highlight a broader role for LLMs as improvers: verification feedback can not only select among existing solutions but also construct stronger candidates beyond the initial pool.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19515v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19515v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] [图像生成] QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training</title>
      <link>https://arxiv.org/abs/2609.19513v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19513v1</guid>
      <pubDate>Thu, 17 Sep 2026 00:04:09 GMT</pubDate>
      <dc:creator>Davide Vitabile, N. Ranjan, Akshay Nambiar et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Davide Vitabile, N. Ranjan, Akshay Nambiar et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, gan, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While major organizations train ever-larger models on private corpora, the open ecosystem lacks STEM-focused synthetic datasets that deliver high per-token learning value efficiently for small models. To address this gap, we introduce QVAC Genesis III, a 191.43B-token, STEM-focused multi-domain synthetic corpus covering 19 domains across several difficulty levels and different educational styles. QVAC Genesis III is built via a dual generation strategy that performs targeted teacher distillation using a weak edge-scale student model as signal: the student&apos;s failures are converted into corrective explanations, while its successes are expanded into contrastive option-level reasoning over all answer choices. We further introduce an LLM-as-a-parser evaluation protocol that extracts final answers from free-form outputs and tracks both accuracy and answer validity. To validate the effectiveness of our QVAC Genesis III data, we conduct controlled from-scratch ablations with 1.7B-parameter models, showing that models trained with QVAC Genesis III consistently outperform both models trained with the open-source synthetic corpus Cosmopedia-v2 and the publicly released Cosmo-1B model across ARC, GPQA Diamond, and MMLU STEM benchmarks, achieving up to +28.57% on ARC-E and +21.35% on ARC-C, while reaching a Valid Answer Rate of up to 99.45%.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19513v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19513v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions</title>
      <link>https://arxiv.org/abs/2609.19512v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19512v1</guid>
      <pubDate>Thu, 17 Sep 2026 00:01:06 GMT</pubDate>
      <dc:creator>Zoe Li</dc:creator>
      <category>模型架构</category>
      <description>Robots can recall prior failures without knowing whether recalled evidence remains valid, conflicts with current observations, or is sufficient to guide a decision. We present CoreSense, a robot-syste...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zoe Li</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-17</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Robots can recall prior failures without knowing whether recalled evidence remains valid, conflicts with current observations, or is sufficient to guide a decision. We present CoreSense, a robot-system integration architecture that combines traceable episodic evidence with a conflict-aware belief gate and bounded, auditable recommendations. The gate checks scope, provenance, time, contradiction, and support before it permits PROCEED, requests re-observation, abstains, or escalates. Evaluation follows three complementary layers without commanding a physical robot: offline public real-robot data, a frozen signal-level simulation, and a live cloud deployment path. On CableTrace-120 and BotFails-200, belief gating reduces protocol-defined unsafe proceeds from 20% and 40% to 0%. A disjointly calibrated raw-video policy also reaches 0% unsafe proceed, but overblocks every nominal episode. On public data, a ViFailback-BotFails visual detector reaches 0.778 AUROC yet remains all-blocking, whereas cycle-disjoint UR3 telemetry for protective stops yields 0% unsafe proceed, 36.1% overblocking, and 61.9% coverage; grip-loss transfer remains a negative result. Controlled physical corroboration yields 3.3%, 0%, and 42.0%, while conflict-aware fusion yields 4.7%, 0%, and 42.8%. Finally, 20/20 cloud recalls validate a CockroachDB Cloud-Amazon Bedrock deployment path. The evidence supports an auditable integration pattern, not autonomous recovery or certified safety.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19512v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19512v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Null importance: Disentangling relevance for interpretable machine learning</title>
      <link>https://arxiv.org/abs/2609.19511v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19511v1</guid>
      <pubDate>Wed, 16 Sep 2026 23:55:03 GMT</pubDate>
      <dc:creator>Garvesh Raskutti, Kris Sankaran, Jiaxin Ye</dc:creator>
      <category>模型架构</category>
      <description>Feature importance is central to interpretable machine learning, but the term &quot;importance&quot; encompasses several fundamentally different notions of relevance. We develop a unified perspective based on n...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Null importance: Disentangling relevance for interpretable machine learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Garvesh Raskutti, Kris Sankaran, Jiaxin Ye</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Feature importance is central to interpretable machine learning, but the term &quot;importance&quot; encompasses several fundamentally different notions of relevance. We develop a unified perspective based on null importance: a population-level characterization of when a feature is irrelevant under a specified notion of relevance. We consider standard notions of null importance arising from marginal and conditional statistical relevance, predictive risk, functional invariance, and causal effects, and show how these notions answer different scientific questions. We illustrate the framework in two applications in which the distinction is particularly consequential: algorithmic fairness, where common fairness criteria correspond to different notions of null importance, and genomic perturbation modeling, where different notions of relevance lead to different conclusions about what a prediction model has learned. The framework connects three aspects of feature analysis: the scientific question defining relevance, the data and model assumptions that shape how different null notions relate, and the methods used to assess importance. We establish sufficient conditions under which null notions coincide and give counterexamples showing how they diverge when those conditions fail. We then characterize which nulls different method families target and when their zero-importance statistics identify those targets. Finally, simulations spanning feature dependence, redundancy, nonlinearity, hidden features and other standard phenomena, along with case studies on image and multiomics data, provide empirical evidence for these theoretical distinctions and their practical consequences. Taken together, these results provide a common statistical language for relating scientific questions, data-generating assumptions, and algorithms, and clarify the conclusions that feature-importance analyses can support.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19511v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19511v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [图像生成] SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features</title>
      <link>https://arxiv.org/abs/2609.19483v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19483v1</guid>
      <pubDate>Wed, 16 Sep 2026 22:51:56 GMT</pubDate>
      <dc:creator>Abdarahmane Traoré, Andy Couturier, Éric Hervet</dc:creator>
      <category>多模态生成</category>
      <category>图像生成</category>
      <description>Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Abdarahmane Traoré, Andy Couturier, Éric Hervet</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can compete. We present SCOUT, which casts cross-modal retrieval as prediction in embedding space. A trainable predictor maps the patch tokens of a frozen video encoder into the embedding space of a frozen text encoder under a bidirectional InfoNCE objective, and no encoder is fine-tuned in the base model. The video encoder is V-JEPA, the text encoder is EmbeddingGemma, and the predictor is initialized from a Qwen3.5-0.8B decoder. We make three findings. First, the best frozen text encoder is simply the one whose geometry best matches the video features. A training-free alignment score ranks three candidate text encoders in the same order as their retrieval accuracy on our held-out split (Spearman $ρ= 1.0$); a fourth, LLM-based encoder shows the rule is metric-dependent, holding for a neighborhood-overlap score ($ρ= 0.8$) but not for a linear probe ($ρ= -0.2$). Second, two precision-targeted levers, parameter-efficient ExPLoRA adaptation of the video encoder and a training-free attribute-decomposed reranker built on a vision-language model, improve the top-rank precision that otherwise limits the frozen system, adding 2.2 points of leaderboard R@1. Third, a local-versus-public calibration study explains which interventions transfer to the real domain. On AI City Challenge 2026 Track 4 the full retrieve-fuse-rerank system reaches 84.25 mAP@10 on the final leaderboard, while a single frozen model submitted alone reaches 60.63. Our trained components cost about 95 GPU-hours. CMP, the dataset authors&apos; fine-tuned cross-encoder that trains for sixteen GPU-days, is one fusion member of the full system, not an alternative. Code and annotations: https://github.com/abtraore/SCOUT-ECCV</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19483v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19483v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization</title>
      <link>https://arxiv.org/abs/2609.19476v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19476v1</guid>
      <pubDate>Wed, 16 Sep 2026 22:39:26 GMT</pubDate>
      <dc:creator>Donney Fan, Colin Doumont, Aleksandra Kalisz et al.</dc:creator>
      <category>图像生成</category>
      <description>Generative models are increasingly central to many de novo discovery pipelines, in which designs are generated at scale and filtered through virtual screens to determine a set of candidates to experim...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Donney Fan, Colin Doumont, Aleksandra Kalisz et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> image generation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Generative models are increasingly central to many de novo discovery pipelines, in which designs are generated at scale and filtered through virtual screens to determine a set of candidates to experimentally validate. While Bayesian optimization (BO) is a natural fit for this setting, as it uses past evaluations to guide future proposals, the computational overhead required for its sequential decision-making becomes a bottleneck when virtual screens are relatively cheap. We make BO practical in this regime by exploiting the unique combination of a linear model constrained to a spherical domain where high-dimensional latents concentrate. We build off recent work justifying the use of linear surrogates, while deriving nearly closed-form solutions to the surrogate modelling and acquisition problems that exploit spherical symmetry. The result is at least a 100x speedup over state-of-the art baselines, with matching or improved performance across molecular and image generation benchmarks. Altogether, our method makes BO a practical drop-in for de novo pipelines where it was previously too slow to consider.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19476v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19476v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Enhanced Agriculture-informed Neural Network by Domain Knowledge</title>
      <link>https://arxiv.org/abs/2609.19466v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19466v1</guid>
      <pubDate>Wed, 16 Sep 2026 22:13:57 GMT</pubDate>
      <dc:creator>Ci Lin, Futong Li, Rose Chong-Wu et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Accurate prediction of nitrous oxide (N2O) emissions from agriculture is important for assessing environmental impacts and supporting sustainable farming. However, prediction remains difficult because...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Enhanced Agriculture-informed Neural Network by Domain Knowledge</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ci Lin, Futong Li, Rose Chong-Wu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Accurate prediction of nitrous oxide (N2O) emissions from agriculture is important for assessing environmental impacts and supporting sustainable farming. However, prediction remains difficult because N2O emissions result from complex interactions among soil properties, climate, biochemical processes, and management practices, while high-quality observations are limited. Deep learning models can capture nonlinear relationships but often lack physical interpretability and may generalize poorly across environmental conditions. We propose the Knowledge-enhanced Agriculture-informed Neural Network (KAINN), a hybrid neural-mechanistic framework that extends the Agriculture-informed Neural Network by incorporating domain knowledge about fertilizer diffusion, soil respiration, and water-filled porosity. We evaluate KAINN using CNN, LSTM, and Transformer architectures across multiple growing seasons and input-feature configurations. The results show that KAINN generally provides lower root mean square error and mean absolute error and higher R-squared values than purely data-driven models and the original AINN. Analysis of the learned interfaces also shows smoother and more physically consistent parameter trajectories with reduced uncertainty. These findings demonstrate that incorporating environmental knowledge into neural networks can improve the reliability, interpretability, and generalization of agricultural N2O-emission predictions.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19466v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19466v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] ParticleSplat: Self-supervised Object-centric Latent Particle Splatting</title>
      <link>https://arxiv.org/abs/2609.19463v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19463v1</guid>
      <pubDate>Wed, 16 Sep 2026 22:05:05 GMT</pubDate>
      <dc:creator>Lyuxing He, Daniel Guo, Elizabeth Terveen et al.</dc:creator>
      <category>模型架构</category>
      <description>We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent &apos;&apos;particles&apos;&apos; representing semantic entities through feedforward 3D Gaussian Spla...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">ParticleSplat: Self-supervised Object-centric Latent Particle Splatting</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Lyuxing He, Daniel Guo, Elizabeth Terveen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent &apos;&apos;particles&apos;&apos; representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which represents images as a set of particles with attributes such as position, scale, and visual appearance, we address a key limitation of DLP: its inherently 2D nature, which prevents explicit 3D spatial and geometric reasoning that are critical for downstream tasks such as robotic manipulation. Leveraging the structural similarity between latent particles and 3D Gaussian primitives, we introduce a 3D latent particle space trained with a novel view synthesis objective. Our model jointly encodes multiple views with camera poses into a shared 3D object-centric latent space, then transforms particles into particle-aligned 3D Gaussians whose composition reconstructs the full scene. On simulated and real-world datasets, we show that this formulation inherently learns object masks without supervision and supports controllable 3D scene editing, such as moving objects by modifying particles in the latent space. We further establish that the learned 3D representation improves downstream performance on robotic manipulation tasks.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19463v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19463v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics</title>
      <link>https://arxiv.org/abs/2609.19453v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19453v1</guid>
      <pubDate>Wed, 16 Sep 2026 21:37:51 GMT</pubDate>
      <dc:creator>Kaitlin Zareno, Jarett Dewbury, Siamak K. Sorooshyari et al.</dc:creator>
      <category>模型架构</category>
      <description>Antimicrobial resistance is expected to claim 10 million lives per year by 2050, and resource-limited regions are most affected. Raman spectroscopy is a novel pathogen diagnostic approach promising ra...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kaitlin Zareno, Jarett Dewbury, Siamak K. Sorooshyari et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Antimicrobial resistance is expected to claim 10 million lives per year by 2050, and resource-limited regions are most affected. Raman spectroscopy is a novel pathogen diagnostic approach promising rapid and portable antibiotic resistance testing within a few hours, compared to days when using gold standard methods. However, current algorithms for Raman spectra analysis 1) are unable to generalize well on limited datasets across diverse patient populations and 2) require increased complexity due to the necessity of non-trivial pre-processing steps, such as feature extraction, which are essential to mitigate the low-quality nature of Raman spectral data. In this work, we address these limitations using Sharpness-Aware Minimization (SAM) to enhance model generalization across a diverse array of hyperparameters in clinical bacterial isolate classification tasks. We demonstrate that SAM achieves accuracy improvements of up to 10.5% on a single split, and an increase in average accuracy of 2.7% across all splits in spectral classification tasks over the traditional optimizer, Adam. These results display the capability of SAM to advance the clinical application of AI-powered Raman spectroscopy tools.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19453v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19453v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] The syntax and semantics of goals</title>
      <link>https://arxiv.org/abs/2609.19448v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19448v1</guid>
      <pubDate>Wed, 16 Sep 2026 21:32:03 GMT</pubDate>
      <dc:creator>David M. Abel, Mark K. Ho</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>In both cognitive science and computer science, goals are conceptualized as cognitive states that flexibly combine with world knowledge to organize and specify purposeful behavior. In this way, goals ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">The syntax and semantics of goals</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> David M. Abel, Mark K. Ho</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">In both cognitive science and computer science, goals are conceptualized as cognitive states that flexibly combine with world knowledge to organize and specify purposeful behavior. In this way, goals are compositional representations whose content relates to rational behavior. We here draw attention to goals as representations and their content because it highlights a parallel with other areas in cognitive science - in particular, the syntax-semantics interface in linguistics and logic - while also foregrounding foundational questions about the expressivity, design, and efficiency of different goal representations. For example, goals are typically taken as fixed and imposing constraints on desirable behaviors, but we can also identify constraints on goal representations themselves, such as whether a particular goal language is sufficiently expressive to capture behaviors of interest, or whether different goal representations capture the same behavior. Here, we synthesize work that aims to characterize the properties of different goal representations and suggest these are points of a broader design space. We close by discussing how distinguishing the form and meaning of goals can elucidate the implicit assumptions we make about goals, inform the study of interactions between higher-level cognition and motivation, and isolate axes of variation for different conceptions of goals.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19448v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19448v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Seeing Abnormal from Normal: Glomerular Abnormality in Representations of Normal Renal Morphology</title>
      <link>https://arxiv.org/abs/2609.19444v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19444v1</guid>
      <pubDate>Wed, 16 Sep 2026 21:27:00 GMT</pubDate>
      <dc:creator>Greta Hasko, Rachit Saluja, Tianyu Shi et al.</dc:creator>
      <category>模型架构</category>
      <description>Fine-grained evaluation of glomerular pathology must distinguish normal glomeruli from abnormalities such as global and segmental glomerulosclerosis, obsolescent, ischemic, solidified, disappearing, a...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Seeing Abnormal from Normal: Glomerular Abnormality in Representations of Normal Renal Morphology</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Greta Hasko, Rachit Saluja, Tianyu Shi et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> u-net</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Fine-grained evaluation of glomerular pathology must distinguish normal glomeruli from abnormalities such as global and segmental glomerulosclerosis, obsolescent, ischemic, solidified, disappearing, and atubular glomeruli. Supervised classification requires labeled examples of every category, which is impractical when subtypes are rare or absent from the training cohort. One-class anomaly detection offers an alternative by modeling normal data and scoring deviations, allowing previously unseen abnormalities to be detected. We use the frozen residual U-Net backbone of Omni-Seg, pretrained to segment structurally normal renal primitives without abnormal-subtype labels. We propose NoRDeC (Normal-Reference Detection and Characterization), a framework combining Mahalanobis normal-reference scoring with layer-wise representation analysis to determine whether and where glomerular pathology is encoded, how spatial aggregation affects detection, and whether abnormalities alter inter-layer relationships differently. Using glomerular images from two institutions, we evaluate backbone layers and aggregation strategies, compare NoRDeC with PaDiM and PatchCore, and analyze representations using centered kernel alignment (CKA). Layer 4 with Center-70 aggregation achieved a pooled AUROC of $0.926\pm0.013$. NoRDeC achieved the highest AUROC in six of seven abnormality categories and in the pooled analysis, while CKA suggested subtype-dependent changes in inter-layer relationships not captured by anomaly scores alone. The normal-reference model is fitted using only normal glomeruli; abnormality labels are used for configuration selection, evaluation, and grouping in the representation analysis. These results show that a frozen renal feature extractor can support both detection and representation-level characterization of glomerular abnormalities without using abnormal examples to fit the detector.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19444v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19444v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Deep Learning Detection of Beyond-General-Relativity Deviations in Gravitational-Wave Signals: A Detection-Threshold Study with Real LIGO Noise</title>
      <link>https://arxiv.org/abs/2609.19416v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19416v1</guid>
      <pubDate>Wed, 16 Sep 2026 20:53:33 GMT</pubDate>
      <dc:creator>Muhammad Adnan Shahzad</dc:creator>
      <category>模型架构</category>
      <description>We study machine-learning detection of controlled beyond-General-Relativity (beyond-GR) deviations in gravitational-wave signals, using both synthetic aLIGO-PSD noise and real LIGO H1 detector strain....</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Deep Learning Detection of Beyond-General-Relativity Deviations in Gravitational-Wave Signals: A Detection-Threshold Study with Real LIGO Noise</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Muhammad Adnan Shahzad</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We study machine-learning detection of controlled beyond-General-Relativity (beyond-GR) deviations in gravitational-wave signals, using both synthetic aLIGO-PSD noise and real LIGO H1 detector strain. Three deviation families are applied to General-Relativistic inspiral-merger-ringdown waveforms: amplitude modulation, phase modulation, and frequency modulation, each parameterized by a dimensionless strength coefficient $β$. A hybrid classifier combining a one-dimensional convolutional neural network with ten hand-crafted waveform statistics is trained on GR and modified waveforms and tested on a deviation type excluded from training. The central result is a quantitative detectability curve as a function of $β$. Using the real GW150914 strain as a template and real H1 detector noise, we find a detection threshold at $β\approx 0.25$, with accuracy rising smoothly from chance at $β\leq 0.2$ to perfect classification at $β\geq 0.5$. The threshold value is specific to the quadratic-in-time modulation form adopted here and should not be interpreted as a generic constraint on beyond-GR parameters. We nevertheless argue that the negative result at small $β$ is informative: it establishes a quantitative limit on machine-learning-only beyond-GR searches in real detector noise, in the absence of matched-filter signal extraction.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19416v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19416v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation</title>
      <link>https://arxiv.org/abs/2609.19414v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19414v1</guid>
      <pubDate>Wed, 16 Sep 2026 20:52:32 GMT</pubDate>
      <dc:creator>Jing Zhang</dc:creator>
      <category>模型架构</category>
      <description>Goal-conditioned reinforcement learning aims to learn policies that reach specified goals, but remains challenging in offline settings with sparse rewards and long-horizon dependencies. In such settin...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jing Zhang</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Goal-conditioned reinforcement learning aims to learn policies that reach specified goals, but remains challenging in offline settings with sparse rewards and long-horizon dependencies. In such settings, goal-completion information can be temporally distant from the early decisions that enable success, while offline value estimation introduces additional error. We study this issue from a reward-propagation perspective and show, in a stylized delayed-goal setting, how goal-directed value separation can become small relative to local estimation error. Motivated by this analysis, we propose Reward Stimulation Implicit Q-Learning (RSIQL), a simple non-hierarchical method that introduces additional reward signals at progress-making intermediate states in offline trajectories. RSIQL uses an auxiliary goal-conditioned value function to identify intermediate states estimated to make progress toward the goal and applies reward stimulation to provide less-delayed training supervision. Unlike hierarchical methods, RSIQL does not learn a separate high-level subgoal policy. Experiments on D4RL goal-reaching benchmarks and OGBench show that RSIQL improves over goal-conditioned IQL on average and achieves performance competitive with hierarchical offline goal-conditioned methods, while retaining a simple flat policy structure.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19414v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19414v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] WZPlanner: Safe End-to-End Path Planning for Autonomous Driving in Work Zones</title>
      <link>https://arxiv.org/abs/2609.19393v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19393v1</guid>
      <pubDate>Wed, 16 Sep 2026 20:22:51 GMT</pubDate>
      <dc:creator>Nishad Sahu, Changzhong Qian, Guangzhou Cai et al.</dc:creator>
      <category>模型架构</category>
      <description>Work zones alter lane geometry through temporary traffic controls and closures that may be absent from on-board maps, challenging autonomous vehicle (AV) perception and planning. Generalization is als...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">WZPlanner: Safe End-to-End Path Planning for Autonomous Driving in Work Zones</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Nishad Sahu, Changzhong Qian, Guangzhou Cai et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Work zones alter lane geometry through temporary traffic controls and closures that may be absent from on-board maps, challenging autonomous vehicle (AV) perception and planning. Generalization is also limited by scarce public datasets with structured geometric supervision. We present WorkZonePlan, a dataset comprising 149K+ synthetic and 5K+ real-world multimodal samples with 3D annotations for lane boundaries, work zone boundaries, and driving trajectory options. It also provides 76 closed-loop CARLA scenarios replayed under three weather conditions, yielding 228 Bench2Drive-format evaluation routes. We introduce WAVE (Work-zone-focused AV data generation in Virtual and rEal Environments), a semi-automated pipeline for creating the dataset, and BoundaryFormer (BF), a transformer-based model that jointly predicts lane and work zone boundary polynomials and driving trajectories. BF uses slot attention for boundary prediction. Ablations show that a separate trajectory decoder using boundary slot features substantially improves trajectory prediction over a slot-attention-only approach. Building on this finding, BF++ offers Camera and Camera+LiDAR variants with metric ground-plane encoding, typed boundary/trajectory queries, long-range point anchors, image-space curve refinement, and conservative gated LiDAR fusion. On the 211 routes common to all four models at the evaluation freeze, BF++-Camera and BF++-Camera+LiDAR achieve Driving Scores of 63.0 and 64.4, respectively, compared with 59.3 for SimLingo and 26.1 for TransFuser++ (TF++). BF++ is 40 times smaller than SimLingo and more than 10 times smaller than TF++, while achieving higher Driving Scores. These results support jointly predicting lane boundaries, work zone boundaries, and driving trajectories as a promising direction toward safer AV operation in work zones. Code and dataset: https://github.com/Nishad-Sahu/WZPlanner.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19393v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19393v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs</title>
      <link>https://arxiv.org/abs/2609.19391v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19391v1</guid>
      <pubDate>Wed, 16 Sep 2026 20:15:36 GMT</pubDate>
      <dc:creator>Albert Wu, Nicholas Roberts, Tzu-Heng Huang et al.</dc:creator>
      <category>模型架构</category>
      <description>LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Albert Wu, Nicholas Roberts, Tzu-Heng Huang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases. Formal verification addresses this by providing machine-checkable guarantees over specified properties, but traditionally demands substantial manual specification and proof engineering. We introduce a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verification-aware intermediate representation where safety properties can be mechanically checked. MAGS formalizes and freezes human-audited APIs and safety requirements, translates generated code into Dafny, repairs violations using verifier feedback, and compiles verified programs back into executable code. We evaluate MAGS on 100 CUDA kernels, 100 terminal scripts, and 20 robotic-arm tasks. Across all 220 examples, it achieves a 100% success rate in producing programs with non-trivial safety guarantees against frozen specifications. Independent safety and functional evaluations further show strong performance across all three domains, while revealing failures when the auto-formalized semantics do not fully capture the target behavior.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19391v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19391v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Do AI Agents Understand Computer Architecture?</title>
      <link>https://arxiv.org/abs/2609.19387v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19387v1</guid>
      <pubDate>Wed, 16 Sep 2026 20:10:28 GMT</pubDate>
      <dc:creator>Ambika Sharan, Grigory Chirkov, Soheil Abbasloo</dc:creator>
      <category>模型架构</category>
      <description>Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Do AI Agents Understand Computer Architecture?</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ambika Sharan, Grigory Chirkov, Soheil Abbasloo</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator may be reasoning about the machine, or may be searching competently over knobs whose meaning it never recovers -- and only the first transfers to the next architecture. Existing evaluations cannot tell the two apart, because they vary the agent while holding the framing of the problem fixed. We do the opposite. AutoTuring hands the same agent the same 15-dimensional accelerator space twice: once as named architectural knobs with simulator counters, once as anonymous variables on [0,1], with the evaluator, the legal space and the reachable optima held identical, so that the only thing that varies is whether the problem means anything. The gap between the two is the measurement. On a nine-kernel FP16 GEMM basket, meaning pays: the architect beats a modeled H200 by 5.4% and its blind counterpart by 12.3% on average, with 70.1% fewer simulator calls. It does not pay uniquely: a critic loop recovers most of that gap for the blind agent and buys the architect nothing, so architectural knowledge and structured critique behave as substitutes rather than as complements. We report these as preliminary findings -- five to six runs per condition on a single modeled accelerator -- and take the comparison itself, not the accelerator, to be the contribution.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19387v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19387v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Mammography Foundation Models for Opportunistic Prediction of Major Adverse Cardiovascular Events</title>
      <link>https://arxiv.org/abs/2609.19385v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19385v1</guid>
      <pubDate>Wed, 16 Sep 2026 20:06:04 GMT</pubDate>
      <dc:creator>Paula Feldman, Nusrat Binta Nizam, Sunwoo Kwak et al.</dc:creator>
      <category>模型架构</category>
      <description>Cardiovascular disease (CVD) remains the leading cause of death among women, yet cardiovascular risk assessment often relies on clinical variables that may be missing, outdated, or unavailable in rout...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Mammography Foundation Models for Opportunistic Prediction of Major Adverse Cardiovascular Events</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Paula Feldman, Nusrat Binta Nizam, Sunwoo Kwak et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Cardiovascular disease (CVD) remains the leading cause of death among women, yet cardiovascular risk assessment often relies on clinical variables that may be missing, outdated, or unavailable in routine care. Screening mammography offers an opportunity for opportunistic cardiovascular risk stratification because it is routinely acquired and contains vascular features, including breast arterial calcifications (BAC), that are associated with cardiovascular risk and events. We evaluate whether mammography specific foundation models, originally pretrained for breast cancer-related tasks, can transfer to cardiovascular risk prediction without cardiovascular specific supervision or explicit BAC annotation. We constructed a 5-year major adverse cardiovascular event (MACE) cohort of 22,497 women linked to electronic health record outcomes, including 500 events (2.22% prevalence). The foundation models achieved AUROCs of 0.823 and 0.822 substantially exceeding an age-only model (AUROC 0.765), despite using only the screening mammogram as input, with no clinical variables. Both foundation models evaluated assigned substantially higher predicted risk to patients with radiologist-documented BAC, despite BAC never being used as a training label, and showed activation patterns consistent with vascular findings. Together, these findings suggest that mammography foundation models can recover clinically relevant cardiovascular risk information directly from mammographic pixels and suggest that screening mammography may provide an opportunistic source of cardiovascular risk information to complement conventional clinical assessment without additional imaging. Code is available in https://github.com/PauFeld/MammoCVD</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19385v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19385v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models</title>
      <link>https://arxiv.org/abs/2609.19384v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19384v1</guid>
      <pubDate>Wed, 16 Sep 2026 20:05:11 GMT</pubDate>
      <dc:creator>Badri N. Patro, Vijay S. Agneeswaran</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offe...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Badri N. Patro, Vijay S. Agneeswaran</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, imagen, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independently trained vision models is difficult when their architectures and parameter shapes differ. Existing weight-space merging methods generally assume aligned, shape-compatible checkpoints, whereas a Vision Transformer (ViT) and a state-space model (SSM) implement token mixing with different operators. We study a hybrid Heterogeneous merging setting that retains both architectures while aligning parameter groups by semantic role. Our proposed Riemannian--Lorentz Parameter Fusion (RLPF) method projects aligned groups to common coordinates, lifts selected coordinates to the Lorentz hyperboloid model of hyperbolic space, computes a regularized geodesic barycenter, and decodes the result into the two branches. A learned gate then combines branch logits for each input. Component groups use fixed curvature values, with normalization parameters treated as Euclidean. In the results available in this manuscript, the fine-tuned system obtains 82.37\% on CIFAR-10, 75.04\% on Oxford-IIIT Pet, and 78.58\% top-1 accuracy on ImageNet-1K; the corresponding best-parent accuracies are 76.54\%, 71.42\%, and 76.42\%. On ImageNet-1K, the reported pre-fine-tuning initialization reaches 77.80\%. These results support further study of geometry-aware heterogeneous fusion, but not a training-free single-checkpoint merge: RLPF is a two-branch hybrid whose gate and reported final models are trained.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19384v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19384v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] FCx: An algorithm for finding Feasible Counterfactual Explanations</title>
      <link>https://arxiv.org/abs/2609.19383v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19383v1</guid>
      <pubDate>Wed, 16 Sep 2026 20:02:57 GMT</pubDate>
      <dc:creator>Kleopatra Markou, Vana Kalogeraki, Dimitrios Gunopulos</dc:creator>
      <category>模型架构</category>
      <description>Counterfactual (CF) explanations identify changes that alter an input&apos;s classification. While existing methods produce realistic and low-cost CFs, they often fail to ensure feasibility, by suggesting ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">FCx: An algorithm for finding Feasible Counterfactual Explanations</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kleopatra Markou, Vana Kalogeraki, Dimitrios Gunopulos</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vae</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Counterfactual (CF) explanations identify changes that alter an input&apos;s classification. While existing methods produce realistic and low-cost CFs, they often fail to ensure feasibility, by suggesting non-constructive modifications or incompatible with future changes (e.g., changing an individual&apos;s race to secure a job offer). We introduce a refinement of CF explanations that explicitly enforces feasibility. Our approach is the first to efficiently generate CFs that are realistic, low-cost and feasible. We accommodate both hard feasible constraints, specified by domain knowledge users, and soft feasible constraints, inferred automatically via causal inference from the dataset. Our method, Feasible Counterfactual Explanations (FCx), is based on a modified Variational Autoencoder (VAE) optimized with a multi-factor loss function. We measure the cost of a change based on the absolute change in values (proximity) as well as the number of features changed (sparsity) while realism is measured based on the LOF for density estimation, guaranteeing that CFs reside in densely populated regions. Extensive experiments on four public datasets show that our approach matches state-of-the-art performance across multiple metrics while guaranteeing feasibility.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19383v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19383v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort</title>
      <link>https://arxiv.org/abs/2609.19374v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19374v1</guid>
      <pubDate>Wed, 16 Sep 2026 19:50:57 GMT</pubDate>
      <dc:creator>Antony Garcia, Gabrielle Britton, Alcibiades Villarreal et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Small clinical tabular datasets require interpretable machine learning because deep learning is often impractical and ensemble models can be difficult to inspect. A key pitfall is that statistical sig...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Antony Garcia, Gabrielle Britton, Alcibiades Villarreal et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Small clinical tabular datasets require interpretable machine learning because deep learning is often impractical and ensemble models can be difficult to inspect. A key pitfall is that statistical significance does not necessarily imply predictive utility. Using data from the Panama Aging Research Initiative--Health Disparities (PARI-HD) cohort (n=165), we implemented a leakage-safe threshold-likelihood Bernoulli/Categorical Naive Bayes (BNB/CNB) classifier. Within every training fold, each continuous predictor was reduced to a supervised chi-square-derived state, while income entered the model through a categorical likelihood. All data-dependent steps were performed within repeated stratified 10-fold cross-validation with 30 repeats. The demographic baseline achieved a ROC-AUC of 0.630 +/- 0.017. I-309 (CCL1) was the dominant incremental feature, increasing AUC by 0.110, with paired DeLong tests yielding p&lt;0.05 in 100% of repeats. In the pre-specified primary analysis, I-309 produced a fixed-partition DeLong p=0.0018, with robustness assessed across 200 random partitions, where the median p-value was 0.0011. Within the exploratory family of 18 candidate markers, I-309 achieved a Benjamini-Hochberg-adjusted q=0.032 on the frozen partition and satisfied q&lt;0.05 in 85% of random partitions, whereas no other marker demonstrated reliable incremental predictive value. Because the fitted model is an inspectable table of thresholds and class-conditional probabilities, these results identify I-309/CCL1 as an interpretable candidate feature for tabular prediction of cognitive impairment, pending external validation.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19374v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19374v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not</title>
      <link>https://arxiv.org/abs/2609.19363v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19363v1</guid>
      <pubDate>Wed, 16 Sep 2026 19:35:47 GMT</pubDate>
      <dc:creator>Rubén Darío Guerrero</dc:creator>
      <category>模型架构</category>
      <description>The query and key projections $\WQ,\WK$ in attention are almost always trained by Euclidean optimizers with no constraint on their geometry. We constrain them to the Stiefel manifold and optimize them...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Rubén Darío Guerrero</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The query and key projections $\WQ,\WK$ in attention are almost always trained by Euclidean optimizers with no constraint on their geometry. We constrain them to the Stiefel manifold and optimize them there with a Riemannian Adam that carries one scalar second moment per frame, caps its step by a trust region, and retracts polarly. Four propositions prove this update is steepest descent in the embedded metric, independent of gradient scale, well conditioned, and exactly $\mathrm{O}(d)$-equivariant, each certified numerically in \texttt{float64}. A fifth supplies the mechanism: weight decay has \emph{identically zero} Riemannian gradient on $\St(d,r)$, since $W = W I_r$ lies in the normal space, so the learned attention geometry survives the collapse cycles that decay drives through the rest of the model. On modular arithmetic grokking, a single run holds $97.0\%$ validation accuracy at epoch 20\,000 against the baseline&apos;s $61.1\%$---an unstable endpoint we report as evidence for the mechanism rather than as an effect size. On CIFAR-10 patches the same rule gains $\mathbf{+8.98}$\,pp over 12 paired starts ($t{=}60.6$, $12/12$), and the gap widens with data rather than eroding. The step rule earns this: a fixed-step Riemannian update is degree one in the gradient, so it moves $24$--$40\times$ less per step than an identically shaped AdamW matrix---its frames barely leave their initialization, and freezing them outright costs only $0.28$\,pp. An ablation credits the whole gain to making the step scale free, and nothing measurable to the projector or to equivariance. A negative result sharpens the account: gauge removal cannot motivate the method, because a direction along which the loss is invariant carries no gradient at all.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19363v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19363v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Smart Insole Human Activity Recognition for Continuous Monitoring in Elderly Care</title>
      <link>https://arxiv.org/abs/2609.19359v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19359v1</guid>
      <pubDate>Wed, 16 Sep 2026 19:31:05 GMT</pubDate>
      <dc:creator>Edwin Rios, Antony Garcia, Fengpei Yuan et al.</dc:creator>
      <category>模型架构</category>
      <description>Falls in older adults are often preceded by changes in mobility, balance, and postural transitions. This paper presents a wireless smart insole platform and machine-learning workflow for recognizing s...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Smart Insole Human Activity Recognition for Continuous Monitoring in Elderly Care</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Edwin Rios, Antony Garcia, Fengpei Yuan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Falls in older adults are often preceded by changes in mobility, balance, and postural transitions. This paper presents a wireless smart insole platform and machine-learning workflow for recognizing sitting, standing, walking, and unstable walking from plantar-pressure and inertial signals. Each insole integrates 16 active pressure-sensing locations and a six-dimensional IMU stream consisting of tri-axial acceleration and angular velocity. Data were collected from 15 healthy adults at 80~Hz and segmented into overlapping windows. Window length and candidate model families were first screened with stratified 10-fold cross-validation; the primary performance estimate was then obtained with participant-independent 5-fold Stratified Group cross-validation, ensuring that all windows from a participant remained in a single fold. Under this protocol, Histogram-Based Gradient Boosting (HGB) achieved macro-F1 scores of 0.954 and 0.959 for the left and right feet, respectively, and 0.980 with bilateral sensing. A compact 1D-CNN evaluated with the same participant-independent folds did not significantly outperform HGB ($p=0.0625$). The results show that low-profile footwear sensing can infer activity state from pressure and IMU measurements for participants unseen during training, establishing a basis for activity monitoring and fall prevention in elderly care.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19359v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19359v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[控制与编辑] [模型架构] [扩散模型] [图像生成] How to Guide Your Language Flow</title>
      <link>https://arxiv.org/abs/2609.19356v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19356v1</guid>
      <pubDate>Wed, 16 Sep 2026 19:29:07 GMT</pubDate>
      <dc:creator>Rohit Dilip, Tianrong Chen, Yuyang Wang et al.</dc:creator>
      <category>控制与编辑</category>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>We introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. This...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">How to Guide Your Language Flow</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Rohit Dilip, Tianrong Chen, Yuyang Wang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 控制与编辑, 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, flow matching, diffusion, conditional generation, diffusion model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. This works using a similar principle as autoguidance, but eliminates the need for an additional forward pass at inference time and provides a reliable path to ensure that the weak and strong model share similar dynamics. We apply and benchmark this method on continuous diffusion language models, where probe guidance sets a new state-of-the-art performance on unconditional generation. When applied to a 1.7B diffusion language model, probe guidance consistently improves on multiple choice question answering benchmarks. Using our probes, we study the traditional autoguidance setting where the strong model is a weak checkpoint, and find that the weak model must come from a low-entropy region of training. These findings both provide a practical way to improve diffusion language models and shed light on the actual mechanism behind autoguidance, which is currently poorly understood.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19356v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19356v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 控制与编辑, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment</title>
      <link>https://arxiv.org/abs/2609.19354v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19354v1</guid>
      <pubDate>Wed, 16 Sep 2026 19:28:43 GMT</pubDate>
      <dc:creator>Henry O. Velesaca, David Freire-Obregon, Luigi Miranda et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the ca...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Henry O. Velesaca, David Freire-Obregon, Luigi Miranda et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the capability of open-source Vision-Language Models (VLMs) to perform zero-shot action quality assessment on Olympic diving videos using the AQA-7 benchmark dataset. In this regard, a regression-based framework is pro-posed to leverage both the semantic reasoning and phase-level sub-scores generated by the VLMs, combining TF-IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. Experimental results show that standalone VLMs achieve moderate Spearman correlations below 0.32, while the proposed ensemble regression framework substantially improves performance in the reported evaluation, reaching a Spearman correlation of 0.67 with a four-model configuration. Textual reasoning features con-sistently outperformed raw numerical sub-scores, highlighting the richness of VLM-generated explanations for action quality analysis. These findings suggest that VLMs hold strong potential as assistive tools for explainable and semi-automated sports performance evaluation. The code is publicly available on GitHub https://github.com/hvelesaca/olympic diving judge vlm</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19354v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19354v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing Process Planning</title>
      <link>https://arxiv.org/abs/2609.19347v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19347v1</guid>
      <pubDate>Wed, 16 Sep 2026 19:22:58 GMT</pubDate>
      <dc:creator>Jingzhan Ge, Ruimin Chen, Azadeh Haghighi et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Robotic additive manufacturing (AM) extends material-extrusion printing beyond gantry kinematics but makes process planning robot-dependent. A slicer-generated plan that appears favorable in part coor...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing Process Planning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jingzhan Ge, Ruimin Chen, Azadeh Haghighi et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Robotic additive manufacturing (AM) extends material-extrusion printing beyond gantry kinematics but makes process planning robot-dependent. A slicer-generated plan that appears favorable in part coordinates can become infeasible or robotically unfavorable on a manipulator because slicer-process decisions and part orientation determine the generated path, while part orientation and workspace placement affect its kinematic realization. Existing AM tools, large language model (LLM)-based decision-support methods, and digital-shadow systems do not provide integrated pre-execution evaluation of these coupled decisions. This paper presents agentic robotic additive manufacturing (A-RAM), an agent-specialist-tool framework that converts user intent and a part file into traceable, execution-ready plans. The LLM interprets manufacturing objectives and constraints, identifies prescribed and searchable planning variables, and encodes this reasoning in a schema-constrained request; a deterministic Planning Agent instantiates the corresponding search workflow, while domain tools compute quantitative evidence for slicing, placement, inverse kinematics, trajectory timing, Joint-6 jerk, and extrusion. The framework is evaluated on a six-axis robotic-arm AM cell through three case studies covering expert-specified planning, goal-only planning, objective-dependent infill screening, and geometry-dependent orientation-placement selection. Across the evaluated candidate sets, selected plans achieve up to 53.5% lower maximum Joint-6 jerk and 48.3% lower mean absolute Joint-6 jerk than the least favorable valid candidates, while objective-specific infill screening yields motion-plan completion times up to 40.1% shorter and extrusion paths up to 12.7% shorter than the corresponding least favorable screened patterns.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19347v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19347v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment</title>
      <link>https://arxiv.org/abs/2609.19325v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19325v1</guid>
      <pubDate>Wed, 16 Sep 2026 18:45:24 GMT</pubDate>
      <dc:creator>Sai Sri Pushpa Jampani, Kshitij Mishra, Asif Ekbal</dc:creator>
      <category>模型架构</category>
      <description>Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithf...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sai Sri Pushpa Jampani, Kshitij Mishra, Asif Ekbal</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remaining hidden from users at deployment. We train this behavior with supervised fine-tuning followed by reinforcement learning with FAITHGATE, a reward-gating objective that grants answer reward only when the safety plan is correct. This discourages safe-looking but unfaithful behavior and promotes tighter plan-answer coupling. Across Qwen backbones, AUDITPLAN improves both robustness and auditability: on Qwen2.5-3B-Instruct, FAITHGATE reduces ASR from 24.0% to 11.6%, LSR from 1.0% to 0.36%, and over-refusal from 11.0% to 2.0%, outperforming answer-only RL, free-form explanation, and weighted-sum structured rewards. Similar trends hold for Qwen2.5-1.5B-Instruct. Larger-model confirmation runs on Qwen-3-4B-Instruct and Qwen2.5-7B-Instruct preserve the same trend suggesting that explicit internal commitments can make safety alignment more faithful, robust, and auditable.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19325v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19325v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning</title>
      <link>https://arxiv.org/abs/2609.19315v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19315v1</guid>
      <pubDate>Wed, 16 Sep 2026 18:25:58 GMT</pubDate>
      <dc:creator>Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta et al.</dc:creator>
      <category>模型架构</category>
      <description>Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason eff...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observability. We present GAVEL, a framework for verifying and repairing long-horizon LLM planning built around an explicit graph world model. The graph represents relevant object-relations, action pre-conditions and effects, and probabilistic beliefs over unobserved object locations. This model can predict the consequences of LLM-generated actions before execution, detect violations, and repair those whose corrections follow directly from the world model. This method also reserves LLM replanning solely for errors requiring semantic reasoning. For multi-task instructions, GAVEL reasons over distributions of possible object locations to reorder remaining subtasks and minimize expected search cost. We evaluate GAVEL on BEHAVIOR-1K across 100 single long-horizon tasks and 500 multi-task instructions. With Qwen3-8B, GAVEL improves single-task success from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%. Distributional belief reasoning also reduces travel distance by approximately 5.4% compared with a static variant. These improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19315v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19315v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Physics-Informed Hemodynamic Modeling for Data-Free Prediction and Sparse-Data Assimilation</title>
      <link>https://arxiv.org/abs/2609.19290v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19290v1</guid>
      <pubDate>Wed, 16 Sep 2026 18:02:24 GMT</pubDate>
      <dc:creator>Xi Chen, Jianchuan Yang, Hongde Li et al.</dc:creator>
      <category>模型架构</category>
      <description>Clinical decision-making for coronary intervention relies mainly on angiography and fractional flow reserve (FFR). However, angiography is two-dimensional and lacks depth information for 3D lesion cha...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Physics-Informed Hemodynamic Modeling for Data-Free Prediction and Sparse-Data Assimilation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xi Chen, Jianchuan Yang, Hongde Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Clinical decision-making for coronary intervention relies mainly on angiography and fractional flow reserve (FFR). However, angiography is two-dimensional and lacks depth information for 3D lesion characterization, while FFR provides only a single functional index, offering limited hemodynamic insight. Among existing methods, numerical analysis is computationally expensive, whereas learning-based approaches require extensive supervision and often lack physical consistency. To address these limitations, we propose physics-informed hemodynamic modeling, an integrated deep learning framework for 3D coronary blood flow analysis from dual-view angiography. First, an attention-enhanced CNN reconstructs coronary geometry from angiography. The resulting point clouds are then mapped to a reference domain and Fourier-encoded for joint representation. A decoupled network separately predicts velocity and pressure fields, with embedded physical priors enabling efficient transfer across physiological conditions. Across 32 clinical patients evaluated under four flow conditions, the trans-stenotic pressure-drop mean absolute percentage error was 2.02%, while the velocity and pressure relative-L2 errors were 0.054 and 0.023, respectively. Validation against hospital-measured FFR further achieved 93.8% diagnostic accuracy (30/32; exact 95% CI, 79.2%-99.2%). The framework also supports illustrative revascularization comparisons and sparse-data assimilation, with the full angiography-to-hemodynamics pipeline completed within 20 minutes per patient.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19290v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19290v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Learning-Induced Dynamical Transition in Recurrent Neural Networks</title>
      <link>https://arxiv.org/abs/2609.19288v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19288v1</guid>
      <pubDate>Wed, 16 Sep 2026 18:01:39 GMT</pubDate>
      <dc:creator>Varun Vaidya</dc:creator>
      <category>模型架构</category>
      <description>Learning in recurrent neural networks can fundamentally reshape their underlying dynamics, transforming initially chaotic activity into stable task-dependent behavior. We develop a non-equilibrium dyn...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Learning-Induced Dynamical Transition in Recurrent Neural Networks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Varun Vaidya</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Learning in recurrent neural networks can fundamentally reshape their underlying dynamics, transforming initially chaotic activity into stable task-dependent behavior. We develop a non-equilibrium dynamical mean-field theory(DMFT) to describe this transition during learning. We show that a slow feedback-driven learning process generates an evolving effective feedback strength that drives the network through a transition from chaotic to stable dynamics defined by a bifurcation of the DMFT solution. By deriving the two-time correlation function throughout learning, we identify a critical feedback strength and a corresponding learning rate dependent critical time separating these regimes. The transition arises from the progressive deformation of an effective dynamical landscape by the growing learned feedback structure. Starting from the untrained state, the theory predicts the time evolution of the network output during training and shows quantitative agreement with numerical simulations.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19288v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19288v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Radio-Frequency Convolutional Neural Networks</title>
      <link>https://arxiv.org/abs/2609.19279v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19279v1</guid>
      <pubDate>Wed, 16 Sep 2026 18:00:07 GMT</pubDate>
      <dc:creator>Zhihui Gao, Shi-Yuan Ma, Yiran Chen et al.</dc:creator>
      <category>图像生成</category>
      <description>Running artificial intelligence (AI) models directly on edge devices such as smartphones, wearables, and drones offers low latency, pervasive scalability, and data privacy, but these devices rarely ca...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Radio-Frequency Convolutional Neural Networks</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhihui Gao, Shi-Yuan Ma, Yiran Chen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> image generation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Running artificial intelligence (AI) models directly on edge devices such as smartphones, wearables, and drones offers low latency, pervasive scalability, and data privacy, but these devices rarely carry the computing capability that modern neural networks demand. Edge accelerators have been developed in response, yet each adds computing hardware to devices already constrained in size, weight, power, and cost (SWaP-C). An alternative lies in what these devices already carry: the frequency mixer in every wireless radio multiplies signals in time, natively performing convolution in the frequency domain. Here we introduce radio-frequency convolutional neural networks (RF-CNNs), which repurpose existing communication hardware for CNN inference. Multi-channel convolutions are mapped onto frequency tones for a passive mixer to execute in a single pass. We experimentally demonstrate that RF-CNN runs deep CNNs up to 26.4 million parameters and nine layers from classification of wireless signals and images to controllable image generation, close to full-precision performance. Because the weights arrive over the air and the analog hardware is shared with communication, the edge device spends energy only on data preparation and readout-down to 0.72 femtojoules per multiply-accumulate, two orders of magnitude less than it would cost on an added digital processor. These results suggest that deployed wireless infrastructure can bring efficient, state-of-the-art AI inference to the billions of devices it already connects.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19279v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19279v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[评估与优化] A Zeroth-Order Paradigm for LLM Preference Alignment</title>
      <link>https://arxiv.org/abs/2609.19144v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19144v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:59:35 GMT</pubDate>
      <dc:creator>Peter Chen, Xi Chen, Wotao Yin et al.</dc:creator>
      <category>评估与优化</category>
      <description>Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Zeroth-Order Paradigm for LLM Preference Alignment</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Peter Chen, Xi Chen, Wotao Yin et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 评估与优化</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> human preference</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we propose and analyze Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles. ComPO extracts directional information from these pairs without directly optimizing a differentiable preference loss on them. We establish a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective. We further introduce online ComPO, which retains the offline comparison mechanism and uses unlabeled policy generations for reverse-KL control relative to a reference policy. Following the coverage perspective of preference fine-tuning, we establish a performance guarantee for a basic constrained scheme under local coverage and in-distribution pairwise reward accuracy. Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 models demonstrate improvements over existing direct alignment methods, including length-controlled win rates, with pair-level diagnostics providing evidence consistent with mitigating likelihood displacement.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19144v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19144v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 评估与优化 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection</title>
      <link>https://arxiv.org/abs/2609.19143v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19143v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:59:30 GMT</pubDate>
      <dc:creator>Sara Pieri, Evangelos Kazakos, Shizhe Chen et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image ca...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sara Pieri, Evangelos Kazakos, Shizhe Chen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or inaccurate segmentation masks. We study this problem through panoptic grounded captioning, a task that requires a VLM to describe both foreground objects and background regions while grounding each referring phrase with pixel-level masks. We make three contributions. First, we introduce PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets. It provides dense captions with near-complete pixel coverage and image-text alignments at the entity level, supporting both training and evaluation. We further propose a phrase-mask matching protocol and a generalized Panoptic Quality (gPQ) metric that jointly evaluates textual and mask agreement. Second, we formulate phrase grounding as selection from a phrase-conditioned pool of mask proposals and introduce PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase. Training this interface jointly with caption generation enables PANORAMA to produce high-quality masks while allowing each phrase to refer to a single region or multiple instances. Third, PANORAMA achieves the best overall grounding on PanoCaps and matches or exceeds specialized models across several pixel-level grounding tasks. Experiments show that our method produces precise entity-level segmentations while maintaining detailed, mask-consistent captions. Code, data and models are available at https://www.di.ens.fr/willow/research/panorama/.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19143v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19143v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics</title>
      <link>https://arxiv.org/abs/2609.19142v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19142v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:59:30 GMT</pubDate>
      <dc:creator>Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung et al.</dc:creator>
      <category>模型架构</category>
      <description>World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into do...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, which excludes web video data from the training pool. We study 3D point track completion as a pre-training objective for learning transferable 3D dynamics without robot data. Given a single RGB-D observation and sparse partial 3D trajectories (tracks), we predict future 3D tracks of all observed points. We show this objective produces a rich 3D dynamics prior, without requiring robot action labels. We contribute a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects, and use it to train PointZero. We show that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data. We demonstrate the utility of our pre-training objective by post-training PointZero for two downstream applications: (1) action-conditioned 3D dynamics prediction and (2) imitation learning. When fine-tuned to condition on end-effector pose, PointZero outperforms the baselines on the recent PGND 3D dynamics benchmark. When fine-tuned to predict robot actions and 3D tracks, PointZero outperforms or matches the baselines on 6/7 simulated and real-world robot manipulation tasks. We furthermore evaluate training PointZero from scratch to isolate the benefits of our proposed architecture from those of our proposed pre-training objective and dataset. We release the dataset, checkpoints, and full training recipe.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19142v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19142v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] In-Context Robot Learning with VLM Agents</title>
      <link>https://arxiv.org/abs/2609.19138v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19138v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:58:35 GMT</pubDate>
      <dc:creator>Dongzhou Cheng, Taoran Yi, Ye Fang et al.</dc:creator>
      <category>多模态生成</category>
      <description>Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">In-Context Robot Learning with VLM Agents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Dongzhou Cheng, Taoran Yi, Ye Fang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19138v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19138v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation</title>
      <link>https://arxiv.org/abs/2609.19137v2</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19137v2</guid>
      <pubDate>Wed, 16 Sep 2026 17:56:44 GMT</pubDate>
      <dc:creator>Guanhua Ji, Tianyu Li, Dayoon Suh et al.</dc:creator>
      <category>视频生成</category>
      <description>Video generation models have advanced rapidly and can now synthesize plausible videos of robot manipulation from image and text prompts. Recent work extracts robot actions directly from such generated...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Guanhua Ji, Tianyu Li, Dayoon Suh et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> video generation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Video generation models have advanced rapidly and can now synthesize plausible videos of robot manipulation from image and text prompts. Recent work extracts robot actions directly from such generated videos, but the result is purely kinematic and lacks force information, causing failures in contact-rich tasks where appropriate contact forces are essential for success. We present a pipeline that jointly leverages generated video and audio to derive motion trajectories and desired-force profiles. The force profile is shaped by the loudness of the generated contact sound, and we execute the resulting force-aware trajectories on a Franka robot using a closed-loop force regulator. We evaluate our pipeline on multiple tasks that require making contact and demonstrate successful zero-shot manipulation where a kinematic-only baseline fails. We also show that the pipeline can be used as a data generation engine to train policies that achieve the tasks in a closed-loop manner. Project website, videos, and dataset: https://dreamingcontactsound.github.io/</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19137v2" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19137v2" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses</title>
      <link>https://arxiv.org/abs/2609.19244v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19244v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:56:32 GMT</pubDate>
      <dc:creator>Mahsa Amani, Seungeon Lee, Abhisek Dash et al.</dc:creator>
      <category>模型架构</category>
      <description>Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversa...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Mahsa Amani, Seungeon Lee, Abhisek Dash et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform&apos;s models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19244v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19244v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging</title>
      <link>https://arxiv.org/abs/2609.19135v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19135v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:56:12 GMT</pubDate>
      <dc:creator>Pranaya Jajoo</dc:creator>
      <category>模型架构</category>
      <description>Can a logged dataset visit every hidden state frequently and still be exponentially uninformative about a target policy&apos;s value? We show that it can when the logger depends on history. For every horiz...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Pranaya Jajoo</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Can a logged dataset visit every hidden state frequently and still be exponentially uninformative about a target policy&apos;s value? We show that it can when the logger depends on history. For every horizon $H \ge 3$, we construct two POMDPs with at most two latent states per stage, three actions, and a common logger with three memory states. Action coverage, belief coverage, and two behavior-marginal outcome-revealing conditions all have constants independent of $H$. Nevertheless, evaluating a known deterministic target policy to accuracy $1/8$ requires $Θ((3/2)^H \log(1/δ))$ logged episodes at confidence $1-δ$, for $0 &lt; δ\le 1/4$, even when both candidate models are known. The mechanism is simple: a reset erases the unknown transition that determines the target value. We characterize the resulting statistical experiment exactly and obtain a matching optimal estimator. A directed two-lane gridworld realizes the construction, and trajectory simulations agree with its finite-sample prediction. The result establishes intractability for the history-dependent-logging, model-based case posed by Zhang and Jiang (2025, arXiv:2503.01134), under their behavior-marginal definition of revealing.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19135v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19135v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Flag Game: A Toy Model for Mechanistic Swarm Interpretability</title>
      <link>https://arxiv.org/abs/2609.19124v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19124v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:46:06 GMT</pubDate>
      <dc:creator>Elizabeth Pavlova, Hidenori Tanaka</dc:creator>
      <category>图像生成</category>
      <description>Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and me...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Flag Game: A Toy Model for Mechanistic Swarm Interpretability</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Elizabeth Pavlova, Hidenori Tanaka</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for collective alignment. To this end, we introduce the Flag Game, a toy model for studying the mechanisms of collective belief formation. Concretely, a hidden country flag defines the ground truth, and each bounded agent directly observes only a private crop but can exchange beliefs and weigh social evidence from peers. Despite its simplicity, the Flag Game reproduces rich collective phenomenology: non-monotonic scaling of performance with population size, accuracy gains from social-awareness prompting and team diversity, and strong effects of organizational structure. In particular, we identify that collective belief collapse at small population sizes turns into collective belief polarization as the population grows. This polarization causes the performance decline at large population sizes, but creates diversity in collective beliefs. Finally, we dissect the mechanisms underlying collective belief collapse and polarization with two complementary approaches. We first introduce social circuit attribution, a technique to predict which agent, and what view, matters most to collective dynamics, and verify its predictions by causal interventions on agents, tracing how agent patching changes collective outcomes. However, the efficacy of causal interventions on agents decreases as the population grows. We therefore develop a statistical mechanical theory for larger populations and verify that it matches the empirical phase diagram. Together, these results take a first step toward mechanistic swarm interpretability, a science of how the properties of individual agents and their communication give rise to emergent collective behavior.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19124v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19124v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation</title>
      <link>https://arxiv.org/abs/2609.19122v2</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19122v2</guid>
      <pubDate>Wed, 16 Sep 2026 17:43:46 GMT</pubDate>
      <dc:creator>Meng&apos;en Qin, Yinchen Liu, Mingxuan Cui et al.</dc:creator>
      <category>图像生成</category>
      <description>Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components wh...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Meng&apos;en Qin, Yinchen Liu, Mingxuan Cui et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> imagen</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components while preserving signal content, but its sparsity coefficient is typically fixed and manually selected. We propose a training-adaptive convolutional sparse coding framework for robust visual signal representation. Specifically, we unfold the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) and treat the sparsity coefficient as a differentiable variable jointly learned with the network parameters. From the information bottleneck perspective, this coefficient controls the trade-off between information retention and compression: the sparsity term promotes compact representations, while the reconstruction term together with task loss preserves task-relevant signal content. We further introduce a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed. Experiments on CIFAR and ImageNet demonstrate competitive clean-data recognition and greatly improved robustness under different input perturbations.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19122v2" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19122v2" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents</title>
      <link>https://arxiv.org/abs/2609.19107v2</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19107v2</guid>
      <pubDate>Wed, 16 Sep 2026 17:36:13 GMT</pubDate>
      <dc:creator>Zixi Chen, Akshay Vegesna, Samip Dahal et al.</dc:creator>
      <category>模型架构</category>
      <description>Scaling laws predict how loss decreases with increases in computation. We show, contrary to conventional wisdom, that architectural interventions can modify scaling exponents in pre-training, leading ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zixi Chen, Akshay Vegesna, Samip Dahal et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Scaling laws predict how loss decreases with increases in computation. We show, contrary to conventional wisdom, that architectural interventions can modify scaling exponents in pre-training, leading to power-law improvements in performance as computation increases. As an anchoring point, we consider the architectural formulation of looped transformers. Although not typically used in this way, looping, also known as recursive depth, provides a mechanism for model growth, by increasing the number of loops during training. Model growth, with and without shared weights, provides the biggest changes to the scaling exponents. In particular, a 7.4B model growth architecture matches GPT-3 13B on CORE with roughly $20\times$ less compute, and has compute efficiency gains that increase with scale. Moreover, simply using a boundary operator in a vanilla transformer, which normalizes and injects an earlier block, also provides an exponent increase, although to a lesser extent. In the data-constrained, multi-epoch setting, standard looping has a useful regularizing effect, where we find it is compute-optimal to increase the number of loops with scale. These results can be understood through the lens of computational depth: for a given computational budget, we wish to increase the usable depth of the transformer, which can lead to efficiency gains that increase with scale.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19107v2" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19107v2" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training</title>
      <link>https://arxiv.org/abs/2609.19242v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19242v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:35:09 GMT</pubDate>
      <dc:creator>Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention commu...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, autoregressive model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory. Conventional context parallelism (CP) shards the combined clean-plus-corrupted sequence by position, communicating shared clean K/V together with block-specific corrupted K/V and their gradients. We observe that the BDLM objective separates over target blocks. We introduce block parallelism (BP), a new distributed parallelism dimension that assigns each corrupted-block computation to one rank. To scale BP to long contexts, we introduce context-sharded block parallelism (CSBP), which also shards the shared clean sequence across those ranks. CSBP keeps corrupted K/V and gradients local, avoids replicated clean prefixes, and preserves BDLM training semantics. On 16 H200 GPUs at 256K context, CSBP improves throughput over the best baseline by 1.18-1.45x for supervised fine-tuning and 1.27-1.33x for conversion of autoregressive models to BDLMs, while matching or reducing peak HBM. Full-model speedup reaches 1.61x at 512K. On eight H100 GPUs, CSBP accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M. In matched 12-hour DiffusionGemma 26B-A4B SFT runs, CSBP achieves higher pass rates at every trained checkpoint on SWE-bench Verified and Terminal-Bench Lite. Code: https://github.com/ScalingIntelligence/Turbo-dLLM</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19242v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19242v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation</title>
      <link>https://arxiv.org/abs/2609.19093v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19093v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:28:51 GMT</pubDate>
      <dc:creator>Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny et al.</dc:creator>
      <category>模型架构</category>
      <description>Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, v...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology, shorthand, formatting, and level of detail. These variations in reporting norms represent an under-appreciated obstacle in efforts to evaluate AI-based radiology report generation (RRG) models, where machine-generated reports are typically assessed based on their concordance with human-generated references. In this paper, we quantify the sensitivity of established evaluation metrics to variations in reporting practices, revealing impacts large enough to alter the rankings of models. We introduce a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of our taxonomy while preserving clinical interpretation. For instance, when comparing the performance of nine RRG models on MIMIC-CXR using RadCliQ-v1, condensing the discussion of normal findings in the reference reports causes Libra to drop from first to second place while CheXOne rises from third to first. Our results suggest that many current metrics fail to decouple clinical interpretation from conformity to reporting practices and that choosing the ``right&apos;&apos; references that accurately reflect the desired reporting practices can be important in practice. To support future research, we release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 (original, alternative) reference report pairs derived from MIMIC-CXR.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19093v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19093v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Securing quantum error correction against misleading advice from AI agents</title>
      <link>https://arxiv.org/abs/2609.19090v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19090v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:26:54 GMT</pubDate>
      <dc:creator>A. Barış Özgüler</dc:creator>
      <category>模型架构</category>
      <description>Can an attacker turn influence over an artificial intelligence (AI) adviser into a harmful quantum error-correction update? We identify an ambiguity in passive syndrome records that obstructs recovery...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Securing quantum error correction against misleading advice from AI agents</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> A. Barış Özgüler</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Can an attacker turn influence over an artificial intelligence (AI) adviser into a harmful quantum error-correction update? We identify an ambiguity in passive syndrome records that obstructs recovery selection, then show how additional calibration measurements support certified recovery updates under uncertainty and drift. In an odd-distance square toric code with error-free preparation, syndrome measurements, and recovery operations, opposite coherent $X$ rotations produce identical passive syndrome-history distributions. Yet a fixed phase correction can help at one sign and harm at the other. A terminal logical measurement on known encoded calibration states supplies the missing sign information. A separate evaluator accepts an update only when calibration uncertainty and a justified drift bound certify improvement over the current recovery, without assuming that the adviser recommends correctly. In simulated advice attacks, calibration-confidence checks reject harmful proposals while retaining beneficial updates under honest advice. We derive sufficient limits on calibration age that require improvement through deployment. In matched simulations, a validated channel-specific bound retains more beneficial updates than the general bound after accounting for evaluation time, while preventing the tested harmful activations under the stated drift assumption. A separate surface-code experiment includes stochastic circuit faults and noise changing during acquisition. Deterministic controllers achieve at least as many beneficial updates with the same observations. Violating the drift assumption permits harmful acceptance in the toric experiment. The results identify information required for recovery selection, establish conditional guarantees against harmful updates, and quantify the recovery improvements forgone through conservative acceptance.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19090v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19090v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education</title>
      <link>https://arxiv.org/abs/2609.19088v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19088v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:26:10 GMT</pubDate>
      <dc:creator>Luyao Zhu, Xun Wei Yee, Wei Li et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language lea...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Luyao Zhu, Xun Wei Yee, Wei Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic imagery, understand its semantic, affective, and cultural content, and reason about visual context to support meaningful interaction. However, existing benchmarks primarily focus on real-world images or domain-specific educational reasoning, providing limited coverage of artistic educational content. To address this gap, we introduce MUSE, a benchmark for evaluating large vision-language models on artistic image understanding in situated educational applications. MUSE decouples image annotation from question generation, enabling diverse tasks with controllable difficulty while reducing annotation effort. It comprises twelve tasks spanning visual perception, semantic and affective interpretation, culture understanding, and compositional reasoning, together with diverse artistic images deliberately curated to center Singaporean and Southeast Asian multicultural contexts alongside Western art traditions, covering multiple themes and difficulty levels. Evaluation of open-source and proprietary models reveals substantial disparities across capability dimensions, particularly in affective interpretation and compositional reasoning. Our analysis further identifies common failure modes and key challenges for developing trustworthy multi-modal models for education. We hope MUSE will serve as a standardized benchmark for advancing multi-modal understanding in situated educational applications.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19088v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19088v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings</title>
      <link>https://arxiv.org/abs/2609.19083v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19083v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:23:50 GMT</pubDate>
      <dc:creator>Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser</dc:creator>
      <category>模型架构</category>
      <description>Kernel methods, and Gaussian Processes (GPs) in particular, require a Hilbertian distance measure---one whose square is conditionally negative definite (CND)---to guarantee positive semi-definiteness ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Kernel methods, and Gaussian Processes (GPs) in particular, require a Hilbertian distance measure---one whose square is conditionally negative definite (CND)---to guarantee positive semi-definiteness (PSD) of the kernel matrix; a condition that fails for many natural input spaces, including smooth manifolds and spaces of probability distributions. We propose the Sparse Landmark Embedding (SLE) kernel, which eliminates this requirement entirely. Each input is embedded into a sparse feature vector via compactly supported bump functions centered at all |D| training points; applying any standard PSD kernel in this embedding space yields a kernel that is provably PSD for arbitrary distance measures. The compact support automatically controls embedding sparsity, keeping kernel matrices well-conditioned and computationally tractable despite the high ambient dimension. We provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and demonstrate, using geodesic and Wasserstein distances, that the SLE kernel matches or substantially exceeds domain-specific baselines in both predictive accuracy and uncertainty quantification.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19083v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19083v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control</title>
      <link>https://arxiv.org/abs/2609.19074v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19074v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:15:59 GMT</pubDate>
      <dc:creator>Bernd Frauenknecht, Emma Cramer, Artur Eisele et al.</dc:creator>
      <category>模型架构</category>
      <description>Reinforcement learning (RL) is an exciting concept as well as a remarkable success story worth sharing. However, RL builds on rather complex interactions between different objects that play out over s...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Bernd Frauenknecht, Emma Cramer, Artur Eisele et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Reinforcement learning (RL) is an exciting concept as well as a remarkable success story worth sharing. However, RL builds on rather complex interactions between different objects that play out over several cycles. Such dynamics are often best explained with an easily accessible implementation. We present RLLBC-Lib, a carefully crafted code library with the goal of lowering the entry barrier for students and other learners of RL in the context of learning-based control. At its heart, RLLBC-Lib comprises a comprehensive library of tabular RL approaches to enforce a clear understanding of the theoretical foundations. A deep RL library follows the same design principles, underscoring the parallels between simple tabular and state-of-the-art deep RL approaches. Additionally, RLLBC-Lib provides a collection of implementations illustrating core RL principles and contrasting RL to other learning-based control approaches. Finally, RLLBC-Lib provides an ideal basis for creating programming assignments with automated grading.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19074v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19074v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment</title>
      <link>https://arxiv.org/abs/2609.19236v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19236v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:05:39 GMT</pubDate>
      <dc:creator>Fangjie Li, Mai Bui, Charan Mohan et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Objective: Incomplete navigation of anatomy during ureteroscopic kidney stone surgeries can contribute to repeat interventions. While skilled surgeons have lower reintervention rates, there are no obj...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Fangjie Li, Mai Bui, Charan Mohan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Objective: Incomplete navigation of anatomy during ureteroscopic kidney stone surgeries can contribute to repeat interventions. While skilled surgeons have lower reintervention rates, there are no objective metrics to quantify scope-navigation performance to evaluate when a trainee becomes skilled. This work aims to recover ureteroscope trajectories from endoscopic video and derive navigation metrics to quantify differences in skill. Methods: We propose RAUL, a reference-assisted reconstruction framework for recovering ureteroscope trajectories from ureteroscope videos only in phantoms. For each phantom, we use a slow, high-quality reference exploration video to generate a reference reconstruction. We localize subsequent exploration videos against this reference. We evaluate localization accuracy against electromagnetically tracked scope pose. We compute navigation metrics from phantom exploration trajectories to compare surgical residents across experience levels. Results: The proposed reference-assisted framework achieves a mean translation root mean square error of $0.5 \pm 0.1$ mm across 9 phantoms. Compared to standard Structure-from-Motion (SfM), the proposed pipeline increases frame-wise localization coverage from $50.5 \pm 14.9\%$ to $86.1 \pm 7.2\%$ of all video frames. The reconstructed trajectories revealed significant differences between high- and low-experience trainees in established navigation metrics. Conclusion: RAUL enables substantially more complete recovery of ureteroscope trajectories from videos compared to standard SfM pipelines, enabling trajectory-based skill assessment without additional tracking equipment. Significance: To the best of our knowledge, this is the first use of video-only recovery of ureteroscope trajectories without external tracking sensors for skill assessment, supporting scalable automated assessment of ureteroscopy navigation skill.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19236v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19236v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Social Laws for Multi-agent Coordination in Stochastic Environments</title>
      <link>https://arxiv.org/abs/2609.18929v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18929v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:05:03 GMT</pubDate>
      <dc:creator>Rolando Fernandez, Caleb Probine, Tyler Lee et al.</dc:creator>
      <category>模型架构</category>
      <description>In multi-agent environments, coordinating agents to prevent interference and ensure robust individual performance is a critical challenge. Previous research on social laws for multi-agent systems has ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Social Laws for Multi-agent Coordination in Stochastic Environments</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Rolando Fernandez, Caleb Probine, Tyler Lee et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">In multi-agent environments, coordinating agents to prevent interference and ensure robust individual performance is a critical challenge. Previous research on social laws for multi-agent systems has primarily focused on deterministic, goal-based settings. This paper extends the concept of social laws to stochastic, reward-based environments, proposing a formalism for defining and verifying their robustness under various conditions. We introduce the notion of $α$-robustness, a measure of the guaranteed utility each agent retains while pursuing its optimal single agent policy, assuming all agents obey the social law. We then present an approach for robustness verification of social laws in stochastic settings, based on a reduction to solving a series of Markov decision processes. Empirical evaluations on toy environments illustrate the potential of our framework.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18929v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18929v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] [图像生成] Comprehensive reconstruction of collider events with hypergraph representation learning and graph-conditioned diffusion</title>
      <link>https://arxiv.org/abs/2609.18928v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18928v1</guid>
      <pubDate>Wed, 16 Sep 2026 17:03:37 GMT</pubDate>
      <dc:creator>Lining Mao, Yvonne Peters, Ethan Simpson et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>In particle collider experiments, event reconstruction is the task of inferring the kinematics of short-lived particles produced in the hard scatter from the stable final states recorded by detectors....</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Comprehensive reconstruction of collider events with hypergraph representation learning and graph-conditioned diffusion</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Lining Mao, Yvonne Peters, Ethan Simpson et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, diffusion model, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">In particle collider experiments, event reconstruction is the task of inferring the kinematics of short-lived particles produced in the hard scatter from the stable final states recorded by detectors. We decompose event reconstruction into two primary tasks: assigning measured jets and charged leptons to parent particles, and predicting unmeasured neutrino kinematics. We present VyPER, a novel geometric learning framework that represents collider events as hypergraphs with a physics-inspired topology. VyPER combines the supervised classification of hyperedges for particle assignment with a diffusion model for predicting neutrino kinematics, leveraging a joint loss function to optimize both reconstruction tasks within a unified framework. We showcase VyPER across several proton-proton collision processes, comparing its performance to existing analytical and machine-learning-based reconstruction techniques. In doing so, we demonstrate that accurate event reconstruction is achievable across a diverse range of Standard Model physics processes, opening new avenues for precision measurements in the Higgs boson, electroweak, and top-quark sectors.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18928v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18928v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image</title>
      <link>https://arxiv.org/abs/2609.18920v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18920v1</guid>
      <pubDate>Wed, 16 Sep 2026 16:58:14 GMT</pubDate>
      <dc:creator>Sneha Paul, Guile Wu, Bingbing Liu et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Physical properties, such as friction, hardness, stiffness, and density, govern how robots should grasp, manipulate and interact with objects, yet estimating these properties from RGB images remains c...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sneha Paul, Guile Wu, Bingbing Liu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, dit, nerf, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Physical properties, such as friction, hardness, stiffness, and density, govern how robots should grasp, manipulate and interact with objects, yet estimating these properties from RGB images remains challenging. Existing methods typically employ per-object reconstruction augmented with physical properties or directly query vision-language models at test time, which results in substantial computational overhead that limits their applicability. In this work, we present PhysVGGT, a feed-forward model that predicts dense maps of friction coefficient, Shore hardness, Young&apos;s modulus, and density, together with object-level mass, from a single RGB image in one forward pass. The key idea of PhysVGGT is to formulate physical property estimation as a dense per-pixel prediction problem and employ a visual geometry transformer to extract geometry-aware tokens from the input image followed by a dense prediction branch for estimating local physical properties and a global prediction branch for estimating object-level mass. In addition, we introduce a scalable pseudo-label generation pipeline that enables large-scale weakly supervised training for dense physical property prediction, substantially reducing the need for expensive direct physical measurements. Extensive experiments show that PhysVGGT achieves state-of-the-art performance on the ABO-500 dataset and generalizes effectively to the out-of-distribution NeRF2Physics dataset. Moreover, PhysVGGT eliminates the need for per-object reconstruction and test-time optimization, achieving an inference latency of only 0.13s per image, making it $27\times$ faster than the previous state of the art.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18920v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18920v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Higher-order pruning of experts in mixture-of-experts language models</title>
      <link>https://arxiv.org/abs/2609.18916v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18916v1</guid>
      <pubDate>Wed, 16 Sep 2026 16:55:46 GMT</pubDate>
      <dc:creator>Alex M. Tseng, Prannay Kaul, Luca Zancato et al.</dc:creator>
      <category>模型架构</category>
      <description>Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count,...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Higher-order pruning of experts in mixture-of-experts language models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Alex M. Tseng, Prannay Kaul, Luca Zancato et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts&apos; contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes an upper bound on the error resulting from pruning. We show that REAP (a state-of-the-art first-order pruning method) is a special case of HOPE where interaction terms are ignored. Across three frontier MoE models (up to 122B parameters), two distinct calibration sets, and multiple benchmarks (including math, instruction following, coding, and an agentic suite), we demonstrate that HOPE produces better pruning decisions than existing methods, and its advantage is most pronounced at high pruning rates and on challenging agentic workloads. At 50% pruning, HOPE outperforms all baselines and achieves an average rank of 1.58 out of 5 methods (versus 2.42 for the next-best method, REAP), with gains of up to +6.1% on agentic coding. Over all conditions, HOPE again achieves the best average rank and surpasses every other method in the majority of head-to-head comparisons. By preserving cooperative expert structure that first-order methods ignore, HOPE enables aggressive compression with minimal degradation, particularly on complex tasks where diverse expert combinations are invoked over long sequences.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18916v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18916v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking</title>
      <link>https://arxiv.org/abs/2609.18909v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18909v1</guid>
      <pubDate>Wed, 16 Sep 2026 16:49:50 GMT</pubDate>
      <dc:creator>Xinshuai Guo, Junjie Wu, Dolly Deng et al.</dc:creator>
      <category>模型架构</category>
      <description>Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in t...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xinshuai Guo, Junjie Wu, Dolly Deng et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> mae</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze large-scale trajectories and identify six complementary process signals that are systematically associated with final agent performance. To disentangle agent performance redundancy from a complete perspective, we propose DualViewEval, an agent benchmark compression method that jointly exploits outcome and process relations to learn an exact-size miniset and predict the full-benchmark scores. Across five agent benchmarks and five representative baselines, DualViewEval achieves the best results in all datasets. With only 20 tasks, it achieves $24\times$--$40\times$ compression on APEX-Agents and BFCL, reducing mean absolute error (MAE) by $14.5\%$--$28.2\%$ over the strongest competitors while improving Kendall&apos;s $τ$ by up to $7.2\%$ relative to EssenceBench on SWE-bench Verified. The selected minisets further reveal capability differences among different agents, providing compact and diagnostic feedback for efficient agentic model development.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18909v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18909v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] NormLift: From Lifted Features To Semantic Reliability In 3D Gaussian Splatting</title>
      <link>https://arxiv.org/abs/2609.18898v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18898v1</guid>
      <pubDate>Wed, 16 Sep 2026 16:35:35 GMT</pubDate>
      <dc:creator>Yihan Zang, Da Li, Dominik Engel et al.</dc:creator>
      <category>模型架构</category>
      <description>Training-free weighted aggregation is widely used to lift 2D semantic features onto 3D Gaussians for open-vocabulary scene understanding, yet its theoretical role remains insufficiently understood. Ex...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">NormLift: From Lifted Features To Semantic Reliability In 3D Gaussian Splatting</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yihan Zang, Da Li, Dominik Engel et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Training-free weighted aggregation is widely used to lift 2D semantic features onto 3D Gaussians for open-vocabulary scene understanding, yet its theoretical role remains insufficiently understood. Existing analyses typically justify this operation from the rendering side, treating Gaussian features as linearly composable Euclidean variables for reconstructing 2D feature maps. However, this view does not match downstream 3D usage, where each Gaussian is often queried independently in a cosine-based embedding space. We revisit feature lifting from the 3D side and formulate per-Gaussian assignment as a cosine alignment problem on the CLIP unit sphere. Under this objective, the L2-normalized semantic back-projected feature emerges as the closed-form solution, providing a complementary interpretation of the standard lifting rule from the perspective of per-Gaussian semantic assignment. The same formulation further yields a norm decomposition into intra-view and inter-view consistency, suggesting that feature magnitude itself can serve as a semantic reliability signal. Calibrated by effective multi-view support, this reliability score guides a mode-voting refinement that preserves CLIP feature validity by avoiding linear averaging. Experiments on open-vocabulary 3D semantic segmentation show that NormLift is an efficient, training-free framework that achieves strong performance across evaluation protocols.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18898v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18898v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Physics-based prediction, uncertainty quantification and decision-making for IN718 crystallographic texture intensity across LPBF defocus regimes</title>
      <link>https://arxiv.org/abs/2609.18863v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18863v1</guid>
      <pubDate>Wed, 16 Sep 2026 16:02:37 GMT</pubDate>
      <dc:creator>Yisheng Lu, John Riris, Jie Song et al.</dc:creator>
      <category>模型架构</category>
      <description>Reliable prediction of crystallographic texture in laser powder bed fusion is critical for linking process conditions with anisotropic response and for qualification. However, black-box models may fai...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Physics-based prediction, uncertainty quantification and decision-making for IN718 crystallographic texture intensity across LPBF defocus regimes</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yisheng Lu, John Riris, Jie Song et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Reliable prediction of crystallographic texture in laser powder bed fusion is critical for linking process conditions with anisotropic response and for qualification. However, black-box models may fail under shift and cannot distinguish weak data support from loss of physical validity. This study develops a two-stage physics-based model for &lt;001&gt; || BD (build direction) texture in Inconel 718. Stage 1 maps process variables to melting mode and melt pool geometry. Stage 2 predicts texture by combining an empirical physics model with a random-forest residual model. A k-nearest-neighbor weight attenuates residual corrections for poorly supported queries, while a study-specific areal beam-power-density criterion withholds predictions outside the adopted conduction envelope. Conformal intervals are evaluated on the retained physics-valid set, and SHAP and Sobol analyses assess residual sensitivity. Under a controlled leave-one-defocus-out evaluation, the physics anchor achieved R^2 = 0.778, against -0.001 for the black-box model and 0.750 for the gated hybrid. Under leave-one-group-out cross-validation, the gated hybrid reached R^2 = 0.592 against 0.538 for the black-box model. Retained-set coverage was 92.9% at a mean full width of 3.65 multiples of a uniform distribution (MUD) under grouped cross-validation and 100% at a width of 3.21 MUD under transfer to a withheld +80 mm defocus regime. An illustrative mapping produced a retained BD elastic-modulus span of 127-187 GPa. On nine conditions from a separately built sample set, the framework withheld three, attenuated three, and matched the measured ordering for the rest. Separating data applicability, physics validity, and predictive uncertainty into distinct decisions lets the framework transfer where an unconstrained model does not, and withhold predictions where no model class performs adequately.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18863v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18863v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] [图像生成] Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection</title>
      <link>https://arxiv.org/abs/2609.18860v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18860v1</guid>
      <pubDate>Wed, 16 Sep 2026 16:00:36 GMT</pubDate>
      <dc:creator>Girish A. Koushik, Diptesh Kanojia, Helen Treharne</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its output. We distinguish these cas...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Girish A. Koushik, Diptesh Kanojia, Helen Treharne</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, lora, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its output. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages $0.740$ versus $0.432$ native macro-F1, while residual reconstruction reaches $0.486$, whereas Gemma improves from $0.532$ to $0.714$. These differences reflect supervised accessibility rather than a pre-existing, native decision rule, and the most influential token role depends on the task. Under the evaluated score scales, Qwen silent-feature ablation is $24-63$ times more probe-sensitive, whereas routed-feature patching on literal yes/no tasks is $16-140$ times more output-sensitive. Calibration-only routing recovers $93.3$% of the mean gap, and probe-distilled LoRA improves native predictions, although shared multi-task adaptation causes negative transfer. A case study of Gemma-3-12B on Facebook Hateful Memes finds a distributed rank-32 image-prompt interaction, reaching $0.756$ versus $0.685$ native macro-F1. Robustness controls show that the signal extends beyond English, is not explained solely by accompanying OCR, and depends on paired visual evidence. Thus, routing, rather than representation alone, is a recurring bottleneck in harmful meme classification.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18860v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18860v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts</title>
      <link>https://arxiv.org/abs/2609.18844v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18844v1</guid>
      <pubDate>Wed, 16 Sep 2026 15:51:51 GMT</pubDate>
      <dc:creator>Liyang Fan, Chi Wei, Yitai Li et al.</dc:creator>
      <category>模型架构</category>
      <description>Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. E...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Liyang Fan, Chi Wei, Yitai Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often isolate short tool calls, API traces, or screenshot resemblance, and a low score under these proxies cannot say whether the model saw poorly, planned poorly, or was failed by its harness. We study scientific overview figure reconstruction, an agent task in which a source image must become an editable PowerPoint slide that preserves text, topology, layout, and native document structure. We introduce ReFigBench, a benchmark and evaluation framework built on 1,000 real overview figures retrieved from arXiv papers with full provenance. Coding agents from four model families reconstruct every figure under two workflows, direct code generation and a specialized PPTX workflow, and the strongest model runs inside two commercial harnesses, yielding ten configurations. Evaluation combines deterministic artifact checks, repeated automated scoring by judges from two model families, and blinded human comparisons. Perception remains a bottleneck that iterative rendering only partly repays. Whether workflow effort converts into quality depends on the model together with its harness, since the same model gains from the specialized workflow inside one harness and loses inside the other, and the harness shifts scores even under an identical direct prompt. The specialized workflow erases native connectors in every configuration, human judges still prefer its renderings in most matchups, and even the strongest agent falls short of the rubric ceiling. These results expose the tension between fidelity and editability as the central challenge for practical multimodal document agents.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18844v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18844v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [模型架构] [扩散模型] Copy What Is Seen, Generate What Is Not: Training-Free Anomaly-Aware Video Restoration</title>
      <link>https://arxiv.org/abs/2609.18836v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18836v1</guid>
      <pubDate>Wed, 16 Sep 2026 15:43:21 GMT</pubDate>
      <dc:creator>Zhida Qu, Shengchao Chen</dc:creator>
      <category>视频生成</category>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>A surveillance system that detects an anomaly often has to repair the footage as well, yet the two tasks are studied in isolation: training-free anomaly detectors stop at a score or a label, while tra...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Copy What Is Seen, Generate What Is Not: Training-Free Anomaly-Aware Video Restoration</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhida Qu, Shengchao Chen</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> video editing, diffusion, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">A surveillance system that detects an anomaly often has to repair the footage as well, yet the two tasks are studied in isolation: training-free anomaly detectors stop at a score or a label, while training-free video editing answers to a user prompt rather than to a detector. This paper proposes AVR (Anomaly-aware Video Restoration), which closes that gap with frozen pretrained models alone and generates content only where the clip offers no evidence to copy. Motion evidence first gates open-vocabulary proposals into spatio-temporal masks. A background prior computed from the clip then fills every pixel the anomaly ever uncovers, leaving diffusion to synthesize only what no frame showed, and a frozen verifier decides per clip whether to trust a classical, a prior-anchored, or a background-conditioned restorer. Extensive experiments on three surveillance datasets, under both full-reference anomaly injection and real anomalies, show that AVR leads full-frame fidelity under oracle masks, matches three trained video inpainters inside the edited region, and outperforms a detect-then-generate pipeline on the masks it produces itself, while suppressing both the residual anomaly and the flicker of free diffusion.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18836v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18836v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings</title>
      <link>https://arxiv.org/abs/2609.19230v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19230v1</guid>
      <pubDate>Wed, 16 Sep 2026 15:34:21 GMT</pubDate>
      <dc:creator>Chao Qin, Fahad Shahbaz Khan, Salman Khan et al.</dc:creator>
      <category>图像生成</category>
      <description>Ultrasound is the most widely deployed imaging modality worldwide, yet clinical AI remains fragmented into narrow single-task models that fail when device, operator, or anatomy changes. Here we presen...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Chao Qin, Fahad Shahbaz Khan, Salman Khan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Ultrasound is the most widely deployed imaging modality worldwide, yet clinical AI remains fragmented into narrow single-task models that fail when device, operator, or anatomy changes. Here we present SonoCorpus, an open resource unifying 456,963 images and 1,626,085 expert masks from 53 public datasets spanning 24 clinical applications and 17 countries, and SonoBase, an interactive segmentation foundation model pretrained on it. Across fifteen evaluation datasets introducing new organs, devices, operators, and geographies, SonoBase outperforms SAM2, MedSAM2, and the concept-promptable MedSAM3 on every dataset and matches per-dataset specialist models trained on the same data; on fully external data it exceeds the accuracy these baselines achieve on their own in-distribution benchmarks. Ejection fraction derived from its segmentations falls within inter-observer variability (6.63\% error), with fewer misclassifications at the defibrillator-candidacy threshold than either promptable baseline (13\% versus 18--42\%); fetal head-circumference (1.81~mm) and gestational-age (1.2 days) errors fall below inter-observer variability. Where a baseline fails outright, one in four test cases, SonoBase recovers a usable segmentation in 81\% of them, including on handheld probes operated by minimally trained users in two low- and middle-income countries (Sierra Leone and Tanzania). Five labeled examples can help the model adapt to a new setting, and the identical training protocol transfers well to newer models such as SAM3, locating the advantage in ultrasound-specific pretraining rather than any single architecture. To ensure reproducibility and enable the community to build on SonoBase as a platform, we release all checkpoints, optimizer states, data-split indices, deduplication hashes, and starter code.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19230v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19230v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] Using OCR Heads to Verbalize Image Semantics</title>
      <link>https://arxiv.org/abs/2609.18823v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18823v1</guid>
      <pubDate>Wed, 16 Sep 2026 15:30:17 GMT</pubDate>
      <dc:creator>Sheridan Feucht, Benno Krojer, Sarah Wang et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Using OCR Heads to Verbalize Image Semantics</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sheridan Feucht, Benno Krojer, Sarah Wang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention heads causally necessary for OCR, and discover that these are in fact general-purpose heads that output interpretable semantic features across all image tokens. For example, pointing these heads at an image token containing the word &quot;bike&quot; causes Qwen3-VL-8B to output &quot;bike,&quot; but pointing them at a bird wing causes the model to output the token &quot;feathers.&quot; We collapse these heads&apos; attention weights into a single verbalization lens transformation that reveals interpretable semantic features in hidden states across all layers. When combined with projection to vocabulary space, we can obtain interpretable labels starting from layer 0, showing that image representations are in fact aligned with language in early layers. We find that we can also use the inverse of this transformation to edit non-word concepts, e.g., replacing a tractor with a revolver in a naturalistic image, providing causal evidence that this subspace is useful for more than just OCR. Our results are an example of how the study of specific mechanisms can shed light on broader interpretability problems.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18823v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18823v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] WaveTLM: Reliable Time-Series Language Modeling through Task Compilation</title>
      <link>https://arxiv.org/abs/2609.18812v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18812v1</guid>
      <pubDate>Wed, 16 Sep 2026 15:21:39 GMT</pubDate>
      <dc:creator>Jiahui Chen, Bingke Zhu, Hongyu Pan et al.</dc:creator>
      <category>模型架构</category>
      <description>Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs. Responses may appear reasonable while halluc...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">WaveTLM: Reliable Time-Series Language Modeling through Task Compilation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Jiahui Chen, Bingke Zhu, Hongyu Pan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs. Responses may appear reasonable while hallucinating the required object: numerical sequences can violate shape, scale, channel order, or temporal alignment, and textual decisions can fall outside the legal label space. We formulate reliable time-series language modeling, separating task-object reliability from predictive quality. We introduce ExecTS-QA, a contract-grounded benchmark spanning forecasting, imputation, classification, anomaly detection, and waveform analysis. We further propose WaveTLM, a unified compiler-executor model whose task compiler transforms user requests, visible arguments, and wave-grounded evidence into typed task states, while task-native executors construct numerical tensors, legal decisions, or structured records. On ExecTS-QA, a single WaveTLM checkpoint achieves 99.40% contract-valid coverage, compared with 37.83% for the strongest evaluated string-first baseline, while retaining balanced predictive performance across all five task families. Evaluations on SciTS, TSQA, IRTS-ToolBench, and ARFBench provide additional evidence of transfer. The code, construction scripts, and ExecTS-QA dataset will be publicly released upon publication. These results show that task compilation can convert plausible language generation into reliable time-series outputs.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18812v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18812v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds</title>
      <link>https://arxiv.org/abs/2609.18782v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18782v1</guid>
      <pubDate>Wed, 16 Sep 2026 15:03:46 GMT</pubDate>
      <dc:creator>Yury Kolomeytsev</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>We establish convergence bounds for deep $V$-learning with horizon $H$. The algorithm fits a scalar value function to targets from executed transitions and selects actions using a predictive model and...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yury Kolomeytsev</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We establish convergence bounds for deep $V$-learning with horizon $H$. The algorithm fits a scalar value function to targets from executed transitions and selects actions using a predictive model and the value function. For current observed-successor targets with fresh true-kernel outcomes, the conditional mean is $\mathcal{T}^βV$, which averages over behavior-policy actions. The Bellman optimality update is $\mathcal{T} V$. We decompose the update error into six residuals: fitting, transition reuse, target construction, replay, action selection, and exploration. Under $L^s$ concentrability, their $L^p$ norms ($p=s/(s-1)$) control expected $L^1$ policy loss. The bound explicitly weights residuals from only the last $H-1$ update blocks, plus an initialization term for shorter runs. We quantify the cost of a shared sampling distribution across horizon levels. For statistical error bounds of order $n^{-ν}$, we derive optimal continuous allocations and an integer allocation whose objective is within a factor $2^ν$ of the constrained optimum. A margin condition with exponent $α$ gives action error of order $Λ^{1+α/p}$, where $Λ$ combines network drift and score error; a one-step construction proves the exponent sharp. Bounds on the distance between frozen and optimal scores transfer an optimal-gap condition to frozen-iterate gap bounds while retaining the mass of optimal ties. Survival probabilities and coverage conditions at deployment yield bounds for policies selected with approximate scores. Separate spatial ReLU networks per horizon level give a conditional neural regression rate, and the finite-state case gives a log-free expected fit rate. These results give expected policy-loss consistency for the fixed-horizon generative-reset approximate-ERM procedure with exact action scores and provide an explicit residual-decay criterion for FIFO/interleaved SGD.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18782v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18782v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] Stable Filters for Generative Modeling of Graph Signals</title>
      <link>https://arxiv.org/abs/2609.18759v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18759v1</guid>
      <pubDate>Wed, 16 Sep 2026 14:46:46 GMT</pubDate>
      <dc:creator>Martin Schmidt, Gonzalo Mateos</dc:creator>
      <category>扩散模型</category>
      <description>Generating signals on graphs requires permutation-equivariant models that exhibit stability with respect to relative structural perturbations. While recent graph-aware Schrödinger bridge models incorp...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Stable Filters for Generative Modeling of Graph Signals</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Martin Schmidt, Gonzalo Mateos</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Generating signals on graphs requires permutation-equivariant models that exhibit stability with respect to relative structural perturbations. While recent graph-aware Schrödinger bridge models incorporate topology information directly into their reference dynamics, it is unclear how perturbations of the graph propagate through these dynamics and affect the resulting generated distributions. In this paper, we analyze the structural stability of graph-aware continuous-time generative models whose drift combines a graph filter with a learned graph neural network. We derive explicit Wasserstein stability bounds that quantify the effect of relative graph perturbations on the generated distributions. Motivated by these bounds, we introduce a principled framework for designing stable graph filters that preserve the smoothing behavior of graph heat diffusion, while boosting structural stability. Experiments on synthetic and fMRI signals show our stable filters enhance structural robustness while matching or exceeding the generative quality of the heat equation baseline.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18759v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18759v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Toward Markerless Video-based Tremor Analysis: Objective Quantification of Pathological Tremor in Mouse Preclinical Models</title>
      <link>https://arxiv.org/abs/2609.18753v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18753v1</guid>
      <pubDate>Wed, 16 Sep 2026 14:44:11 GMT</pubDate>
      <dc:creator>Yota Koshimoto, Akihiro Tsukahara, Yasuhiro Moriwaki et al.</dc:creator>
      <category>模型架构</category>
      <description>Tremor is a movement disorder characterized by involuntary, rhythmic oscillations of body parts and is a hallmark of several neurological conditions, including Parkinson&apos;s disease and essential tremor...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Toward Markerless Video-based Tremor Analysis: Objective Quantification of Pathological Tremor in Mouse Preclinical Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yota Koshimoto, Akihiro Tsukahara, Yasuhiro Moriwaki et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Tremor is a movement disorder characterized by involuntary, rhythmic oscillations of body parts and is a hallmark of several neurological conditions, including Parkinson&apos;s disease and essential tremor. Elucidating its underlying mechanisms relies heavily on mouse models, which offer genetic manipulability and translational relevance to human neural circuitry. Accordingly, these models are indispensable for studying tremor pathophysiology. So far, electromyography and accelerometers have been used as methods to quantitatively observe tremors in mice. However, these methods have several drawbacks, such as high costs and complex setups. In particular, the invasive surgical implantation of devices causes significant stress to the animals. Although RGB-based methods offer non-invasive and cost-effective alternatives, they often lack the sensitivity required to detect subtle tremors. Therefore, this paper addresses these challenges by achieving mouse tremor severity estimation using conventional RGB cameras only. To address the challenging task of isolating tremor-related vibrations while the mouse itself is also in motion, our pipeline incorporates segmentation-based pre-processing to extract the mouse region and a Tremor Score Estimation Module that captures subtle tremors with high sensitivity. In the experiments, we assessed tremors in unrestrained mice using a non-invasive method with two standard cameras. The results demonstrated a strong correlation with accelerometer measurements and confirmed that the method accurately captured the intensity-dependent characteristics of tremors. The project page is available at https://isogawalab.github.io/Video-based-Tremor-Analysis-Project/.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18753v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18753v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] When Edit Flows are Edit Jumps: replicating Edit Flows and EvoFlows</title>
      <link>https://arxiv.org/abs/2609.18745v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18745v1</guid>
      <pubDate>Wed, 16 Sep 2026 14:39:26 GMT</pubDate>
      <dc:creator>Gabriel Bénédict, Melanie Buechler, Gerard Riera-Solà et al.</dc:creator>
      <category>模型架构</category>
      <description>Antibody lead optimization calls for a small, bounded set of edits to an existing candidate: substitutions, but also insertions and deletions. Edit-based generative models are the only ones that alloc...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">When Edit Flows are Edit Jumps: replicating Edit Flows and EvoFlows</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Gabriel Bénédict, Melanie Buechler, Gerard Riera-Solà et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Antibody lead optimization calls for a small, bounded set of edits to an existing candidate: substitutions, but also insertions and deletions. Edit-based generative models are the only ones that allocate such an edit budget without fixing the edit positions, the edit count, or the output length in advance. However, the existing approaches Edit Flows and EvoFlows did not release code or complete training specifications. Here, we show that both methods follow the same underlying process -- edits firing one at a time, at learned rates, in continuous time -- the pure-jump case of generator matching over finite sequences. With EditJumps we introduce the first open implementation of this framework, with a single generalist antibody editor trained on 1.66M Observed Antibody Space homolog pairs to propose homolog-like variants of a seed sequence, editing unseen leads zero-shot, without the per-family retraining original approaches require. Replicating this system from scratch exposes why open code is essential for generative biology: reconciling published edit distributions required reverse-engineering an undocumented rate-scaling hyperparameter that dictates realized mutation counts. Moreover, we show that published evaluation metrics are highly sensitive to reference sample size, frequently flipping method rankings. We release our full codebase, automated test suite, and configurations at: https://github.com/VisiumCH/editjumps</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18745v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18745v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Mask IPL: Noise-Free Intrinsic Position Learning via Computation Graph Clipping for Event-Based Spike-Driven Tracking</title>
      <link>https://arxiv.org/abs/2609.18716v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18716v1</guid>
      <pubDate>Wed, 16 Sep 2026 14:21:42 GMT</pubDate>
      <dc:creator>Yimeng Shan, Malu Zhang</dc:creator>
      <category>模型架构</category>
      <description>Spiking Neural Networks (SNNs) match the event-driven nature of event cameras and naturally extract spatiotemporal features. These properties have motivated a series of recent studies on event-based t...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Mask IPL: Noise-Free Intrinsic Position Learning via Computation Graph Clipping for Event-Based Spike-Driven Tracking</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yimeng Shan, Malu Zhang</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Spiking Neural Networks (SNNs) match the event-driven nature of event cameras and naturally extract spatiotemporal features. These properties have motivated a series of recent studies on event-based tracking with SNNs. Intrinsic Position Learning (IPL) acquires strong position information without introducing additional parameters, making it a mainstream approach for position encoding in event-based spike-driven tracking. However, the mechanism behind its effectiveness lacks systematic theoretical analysis. Moreover, our analysis reveals that IPL introduces noise in both forward and backward propagation. The former increases inference error, while the latter prevents parameters from converging to better solutions. This paper presents a systematic analysis of IPL and demonstrates that its effectiveness stems from the synergy between IPL and multi-stage convolution. The zero blocks in the joint tensor act as zero padding for convolution, and the resulting boundary effect propagates layer by layer through multi-stage convolution. Every parameter update is therefore driven by a gradient that perceives the relative displacement between template and search frames. Positional encoding added after the convolutional stage cannot provide this information. We further propose a simple Computation Graph Clipping method that applies a validity mask determined by the layout to the operations of every layer, making invalid regions equivalent to zero padding in both forward and backward propagation. This eliminates the noise without introducing additional parameters and makes the actual gradient coincide with the ideal gradient. We name the improved method Mask IPL. Without increasing parameters or computational cost, Mask IPL improves the AUC of the Tiny-scale tracker on FE108, FELT, and VisEvent, and consistently improves the Base-scale tracker as well.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18716v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18716v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection</title>
      <link>https://arxiv.org/abs/2609.18690v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18690v1</guid>
      <pubDate>Wed, 16 Sep 2026 14:06:38 GMT</pubDate>
      <dc:creator>Liyang Fan, Xinping Bi, Yitai Li et al.</dc:creator>
      <category>多模态生成</category>
      <description>Graphical User Interface (GUI) grounding is a fundamental perception task for multimodal agents, enabling them to interpret natural language instructions and interact with digital interfaces. Existing...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Liyang Fan, Xinping Bi, Yitai Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Graphical User Interface (GUI) grounding is a fundamental perception task for multimodal agents, enabling them to interpret natural language instructions and interact with digital interfaces. Existing methods face a fundamental trade-off between accuracy and efficiency: direct full-image inference often fails to capture small or visually similar UI elements, while multi-crop strategies improve localization at the cost of multiple expensive Vision-Language Model (VLM) calls per query.   To address this challenge, we propose RankGround, a two-stage framework that achieves accurate GUI grounding with a single VLM call per query. Central to our approach is GroundRanker, a lightweight multimodal reranker that identifies the most promising crop from a dense candidate set. Because no off-the-shelf ranking dataset is available, we construct ranking supervision data from existing grounding datasets. A strict containment criterion and boundary-aware positive augmentation improve alignment and spatial coverage in cluttered layouts. GroundRanker is then trained with a two-stage curriculum: a pointwise objective first learns coarse containment, and a listwise objective refines subtle semantic and spatial distinctions among visually similar crops.   Experimental results show that RankGround consistently outperforms strong baselines while reducing computational cost. It achieves 1.4 times faster inference and improves localization accuracy by 5.5% on average over the second-best method across all backbones and screen scales, establishing a new state of the art in both efficiency and precision for GUI grounding.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18690v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18690v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging</title>
      <link>https://arxiv.org/abs/2609.18688v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18688v1</guid>
      <pubDate>Wed, 16 Sep 2026 14:04:24 GMT</pubDate>
      <dc:creator>Johannes Kaiser, Florian Braunmiller, Daniel Rückert et al.</dc:creator>
      <category>图像生成</category>
      <description>AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (MoE) architectures currently f...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Johannes Kaiser, Florian Braunmiller, Daniel Rückert et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> imagen</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (MoE) architectures currently fail to satisfy. Expert-based routing improves in-domain learning but may sacrifice cross-modal signals, which appear particularly important for rare (low-prevalence) pathologies in our experiments. To resolve this, we introduce Generalist-Specialist-MoE (GS-MoE), a two-branch (MoE) architecture that couples a cross-modal generalist model with distinct modality-specific specialists (experts) via domain-constrained feature fusion. On RadImageNet (1.35M images, 165 pathologies, three modalities), GS-MoE recovers detection of six low-prevalence pathologies on which every baseline scores F1 $=$ 0, with per-class gains up to +0.60 F1. It attains this while even slightly exceeding dense and specialist-only MoE aggregate baselines (MCC 0.770), while using ${\sim}53\%$ fewer active parameters at inference than the strongest investigated dense model.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18688v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18688v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Learning to Program Adaptive Non-Local Observables for Machine Learning</title>
      <link>https://arxiv.org/abs/2609.18655v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18655v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:35:31 GMT</pubDate>
      <dc:creator>Yu-Ting Lee, Samuel Yen-Chi Chen, Huan-Hsin Tseng</dc:creator>
      <category>模型架构</category>
      <description>Quantum neural networks (QNNs) are typically built from variational quantum circuits (VQCs), which are limited by local measurements. Adaptive non-local observables (ANO) address this by jointly optim...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Learning to Program Adaptive Non-Local Observables for Machine Learning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yu-Ting Lee, Samuel Yen-Chi Chen, Huan-Hsin Tseng</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Quantum neural networks (QNNs) are typically built from variational quantum circuits (VQCs), which are limited by local measurements. Adaptive non-local observables (ANO) address this by jointly optimizing circuit parameters and multi-qubit measurements. However, existing ANO-based VQCs learn only a single static observable that remains invariant across all inputs. We propose QFWP-ANO, a novel architecture which employs a classical hypernetwork to dynamically program VQC parameters and/or non-local observables conditioned on each input. On multivariate time-series forecasting across four ETT datasets, QFWP-ANO achieves the lowest MSE in 16 of 20 settings and second-lowest in the remaining four, surpassing ANO-based and other strong baselines. On reinforcement learning tasks, QFWP-ANO consistently surpasses ANO-VQCs. Our results establish input-conditioned ANO as an effective approach for enhancing QNNs.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18655v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18655v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Fallacy Benchmarks Measure Scheme Recognition, Not Fallacy Detection</title>
      <link>https://arxiv.org/abs/2609.18644v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18644v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:28:39 GMT</pubDate>
      <dc:creator>Navyansh Singh, Animesh Pathak, Aarav Singh</dc:creator>
      <category>模型架构</category>
      <description>Fallacy-detection benchmarks pair fallacy classes with a single &quot;valid&quot; or &quot;none&quot; class that takes everything data collection did not label as a fallacy. This construction is misleading: a classifier ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Fallacy Benchmarks Measure Scheme Recognition, Not Fallacy Detection</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Navyansh Singh, Animesh Pathak, Aarav Singh</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Fallacy-detection benchmarks pair fallacy classes with a single &quot;valid&quot; or &quot;none&quot; class that takes everything data collection did not label as a fallacy. This construction is misleading: a classifier can learn cues that do well on this class without learning to tell a fallacy from a correct argument. We show that the low false-positive rates benchmarks report are an artifact of how the class is built, not evidence of detection ability. The most informative negative for a fallacy is a correct argument using the same argumentation scheme, and such arguments are at most a few percent of the valid class across the four benchmarks we examined. Evaluated on constructed scheme-matched negatives, false-positive rates rise from 16.6% to 58.9% on CoCoLoFa and from 5.7% to 62.0% on Reddit. That rate depends on how the negatives are written, so we also compare two conditions from the same pipeline that differ only in scheme identity. Classifiers label scheme-matched negatives as the source fallacy type 40.9 points more often than wrong-scheme negatives, which are instead identified as the scheme they actually use 85.9% of the time against 0.4% for the source type. The classifier has learned which scheme an argument uses, not whether it uses it correctly, and on the benchmarks&apos; own test sets the two are indistinguishable. The same dissociation appears in three zero-shot LLM detectors that never saw these benchmarks, and the measurement is far lower on a negative class that was built deliberately. We release the items as Scheme Foils. A reported false-positive rate should not be trusted as a measure of detection until the valid class has been audited for scheme-matched coverage.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18644v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18644v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] VibeAvatar: Aligning Phonetic Kinematics and Human Aesthetics for High-Fidelity Talking Avatar Synthesis</title>
      <link>https://arxiv.org/abs/2609.18632v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18632v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:20:07 GMT</pubDate>
      <dc:creator>Qilin Wang, Mingyu Li, Hao Tang</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusion-based methods, existing approaches still strugg...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">VibeAvatar: Aligning Phonetic Kinematics and Human Aesthetics for High-Fidelity Talking Avatar Synthesis</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Qilin Wang, Mingyu Li, Hao Tang</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusion-based methods, existing approaches still struggle to jointly achieve accurate lip articulation, human-preferred motion aesthetics, and efficient inference. We observe that phonetic accuracy and motion aesthetics arise from fundamentally different sources and should be addressed at complementary stages rather than learned implicitly by a single generator. Based on this insight, we propose VibeAvatar, which disentangles these two objectives through a Phonetic Kinematics Adapter (PKA) that converts recognition-oriented speech features into phonetic-kinematic conditions at the conditioning stage, and an Aesthetic Motion Policy (AMP) that optimizes a flow-consistent stochastic sampling policy via Group Relative Policy Optimization (GRPO) at the post-training stage. With a lightweight flow-based motion generator operating in a compact 1D warp-based latent motion space, VibeAvatar achieves state-of-the-art results in articulation, aesthetics, and efficiency on both objective metrics and user studies, while generating a 10-second 512px video in under 10 seconds with only $\sim$3GB VRAM.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18632v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18632v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] MoQSplat: Adaptive Progressive Streaming of 3D Gaussian Splatting via MoQ</title>
      <link>https://arxiv.org/abs/2609.18624v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18624v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:14:01 GMT</pubDate>
      <dc:creator>Emanuele Artioli, Mohammadreza Ghafari, Md Tariqul Islam et al.</dc:creator>
      <category>模型架构</category>
      <description>3D Gaussian Splatting (3DGS) enables photorealistic novel view synthesis, but transmitting gigabyte-scale scene data remains challenging for immersive applications. Traditional HTTP Adaptive Streaming...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MoQSplat: Adaptive Progressive Streaming of 3D Gaussian Splatting via MoQ</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Emanuele Artioli, Mohammadreza Ghafari, Md Tariqul Islam et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">3D Gaussian Splatting (3DGS) enables photorealistic novel view synthesis, but transmitting gigabyte-scale scene data remains challenging for immersive applications. Traditional HTTP Adaptive Streaming over TCP introduces Head-of-Line (HOL) blocking and coarse segmenting ill-suited to fine-grained 3DGS delivery. We propose MoQSplat, which maps 3DGS content onto the Media over QUIC (MoQ) transport hierarchy. MoQSplat partitions scenes into spatial Tracks, clusters splats into semantically coherent Groups, and constructs progressive-quality Subgroups mapped to independent QUIC streams to eliminate connection-level HOL blocking. Using a stateless, subscriber-driven adaptation loop, clients dynamically request spatial regions and quality tiers based on six degrees of freedom (6-DoF) frustum visibility, distance, and foveal alignment. We evaluate the core components on a prototype implementation, showing that opacity-based pruning outperforms scale-based pruning for progressive delivery. The source code is available at https://github.com/emanuele-artioli/MoQSplat.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18624v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18624v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory</title>
      <link>https://arxiv.org/abs/2609.18623v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18623v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:13:18 GMT</pubDate>
      <dc:creator>Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao et al.</dc:creator>
      <category>模型架构</category>
      <description>State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal me...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory (RAM), a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes $\sim$10% more routes without traffic rule infractions than the previous state-of-the-art VLA on the challenging Bench2Drive closed-loop driving benchmark. Non-reactive open-loop simulation on the large-scale real-world NVIDIA Physical AI AV dataset shows 10.2% and 7.7% lower collision-violation rates than SimLingo in single- and four-view settings, respectively. Additionally, FIVE-VLA runs at $\sim$30 fps on an A100 and $\sim$4 fps on a T4 GPU (proxy to an edge device), representing an 8-30$\times$ speedup over previous methods.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18623v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18623v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] How Many Labels Does Model Choice Need? Certificates and Budgets for Selective Prediction</title>
      <link>https://arxiv.org/abs/2609.18622v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18622v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:10:59 GMT</pubDate>
      <dc:creator>Tetsuji Kuboyama</dc:creator>
      <category>模型架构</category>
      <description>Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently. We quantify this requirement for the area und...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">How Many Labels Does Model Choice Need? Certificates and Budgets for Selective Prediction</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tetsuji Kuboyama</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently. We quantify this requirement for the area under the generalized risk-coverage curve (AUGRC). A prelabel lower bound rules out insufficient budgets. With all labels known, a covering linear program bounds the minimum number of labels sufficient to fix the winner (the certificate size) within $K-1$ labels for $K$ candidates. For fixed $K$, independent uniform orders and identical predictions, the prelabel bound approaches one quarter of the pool. With iid Bernoulli errors independent of the orders, every exact acquisition policy reads almost all labels asymptotically, although a two-candidate certificate needs only half. Across 108 feature-panel comparisons on nine datasets, disagreement labels settle every accuracy choice but no AUGRC choice. A 20% budget is ruled out in 96 conditions; certificates need 56-57% on average. On ten conditions with pretrained image classifiers, confidence-score choice reads 68-91% of 10,000 labels for exact selection and 50-67% with AUGRC tolerance $5\times10^{-4}$. An exact stopping test works with any acquisition order. Together, these results link confidence ranks to label budgets and certified model comparison.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18622v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18622v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Learning Array Signal Topologies as Conditional Neural Manifolds</title>
      <link>https://arxiv.org/abs/2609.18616v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18616v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:04:28 GMT</pubDate>
      <dc:creator>Julian P. Merkofer, Vincent van de Schaft, Ruud J. G. van Sloun</dc:creator>
      <category>模型架构</category>
      <description>Subspace methods such as multiple signal classification (MUSIC) achieve super-resolution direction of arrival (DoA) estimation by exploiting the orthogonality between the array manifold and the noise ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Learning Array Signal Topologies as Conditional Neural Manifolds</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Julian P. Merkofer, Vincent van de Schaft, Ruud J. G. van Sloun</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Subspace methods such as multiple signal classification (MUSIC) achieve super-resolution direction of arrival (DoA) estimation by exploiting the orthogonality between the array manifold and the noise subspace of the measurements. Their accuracy therefore depends on the assumed manifold and degrades under model mismatch, while parameters not identifiable from the spatial manifold cannot be recovered. In this work, we propose the conditional neural manifold (CNM), which replaces the fixed manifold with an observation-conditioned mapping from source parameters to steering vectors. An encoder maps the snapshots to a latent scene representation that conditions a zero-initialized neural field over the parameter space. The manifold is learned without steering-vector supervision by shaping the resulting MUSIC landscape. Since the correction acts on the manifold rather than on the estimator, it can be used by other manifold-based methods without modification. The CNM restores resolution under array imperfections, colored noise, correlated sources, and near-field propagation, and resolves the angle-frequency ambiguity inherent to the nominal spatial manifold.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18616v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18616v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence</title>
      <link>https://arxiv.org/abs/2609.18612v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18612v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:03:12 GMT</pubDate>
      <dc:creator>Sebastian Gerstner, Hilal AlQuabeh, Kentaro Inui et al.</dc:creator>
      <category>模型架构</category>
      <description>We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs). We propose a simple analysis method: For each neuron, we compute the cosine similarities between its ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sebastian Gerstner, Hilal AlQuabeh, Kentaro Inui et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs). We propose a simple analysis method: For each neuron, we compute the cosine similarities between its input (reading) and output (writing) weight vectors. In this scheme, a strong negative cosine similarity indicates the neuron weakens the direction it detects in the residual stream, so we call this a weakening neuron. This allows us to gain a number of novel insights. First, we show that nine different LLMs have similar patterns: weakening neurons appear mostly in late layers whereas their counterparts, (conditional) strengthening neurons, are frequent in early-middle layers. Second, we find that weakening neurons display surprising behavior: even though there are few, they activate often and have a large influence on model behavior. Third, weakening neurons have a strong effect on model output when gate values are negative -- which is surprising since negative gate values are not expected to encode functionality.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18612v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18612v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A Geometric Theory of Decision Boundaries in Structured Markov Decision Processes</title>
      <link>https://arxiv.org/abs/2609.18610v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18610v1</guid>
      <pubDate>Wed, 16 Sep 2026 13:01:13 GMT</pubDate>
      <dc:creator>Fredy Pokou</dc:creator>
      <category>模型架构</category>
      <description>Classical dynamic programming represents optimal sequential decisions through value functions and policies. While this functional representation is natural for computing optimal decisions, it does not...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Geometric Theory of Decision Boundaries in Structured Markov Decision Processes</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Fredy Pokou</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Classical dynamic programming represents optimal sequential decisions through value functions and policies. While this functional representation is natural for computing optimal decisions, it does not directly identify the mathematical object governing policy reconstruction, representation complexity, or oracle-query complexity once an optimal policy is fixed. This paper addresses this question by developing a geometric theory of structured optimal policies in which the decision-boundary geometry induced by the policy becomes the primary object of analysis. We show that, under suitable structural regularity conditions, this geometry provides the minimal representation required for policy reconstruction and determines the statistical and computational complexity of the reconstruction problem. Building upon this representation, we establish structural properties of policy-induced decision geometry, introduce intrinsic notions of boundary and decision complexity, derive information-theoretic measures of decision compression, and obtain statistical guarantees for boundary estimation and policy reconstruction from black-box policy queries. Collectively, these results demonstrate that, for the structured decision problems considered here, the complexity of policy reconstruction is governed by the geometry of the decision boundary rather than by the cardinality of the ambient state space. Controlled numerical experiments examine the principal theoretical predictions and provide empirical evidence consistent with the proposed framework.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18610v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18610v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?</title>
      <link>https://arxiv.org/abs/2609.18605v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18605v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:57:15 GMT</pubDate>
      <dc:creator>Mika Okamoto, Ansel Kaplan Erol</dc:creator>
      <category>模型架构</category>
      <description>As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules spe...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Mika Okamoto, Ansel Kaplan Erol</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules specified in an agent&apos;s system context is a first-order legal concern. Currently, no evaluation framework systematically measures which LLM models tend to violate compliance rules, especially under pressure from a persistent user, a hurried manager, or circumstances where violation is convenient or attractive. We introduce PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following under pressure in AI agents assisting employees in daily tasks across twelve regulated enterprise domains and forty-eight scenarios, each set in a realistic multi-turn conversation. Each benchmark item pairs a standing rule against a rule-violating shortcut, and applies a battery of pressures across different wordings and system-prompt modes. We construct PACT component by component under strict LLM-as-judge auditing to ensure samples are unambiguous, ungameable, and realistic enough to avoid eliciting evaluation-aware behavior. We use PACT to profile LLM compliance across six complementary metrics that create a holistic picture of an AI assistant&apos;s robustness under pressure and throughout multi-turn conversations, its transparency, and ability to correctly discern where a rule applies. We aggregate this profile into PACTScore, a reliability-weighted compliance rate over all items and modes. Our results across 22 common LLM models spanning multiple providers and sizes show substantial variability in compliance across models and metric dimensions. Even the strongest assistants mis-apply a rule on 6 to 10% of items, and ordinary user pressure raises the violation rate by 65% on average. PACT highlights compliance risks in LLM assistants, motivating guardrails and careful model selection.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18605v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18605v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Online Robust Reinforcement Learning Through Monte-Carlo Planning</title>
      <link>https://arxiv.org/abs/2609.18599v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18599v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:52:13 GMT</pubDate>
      <dc:creator>Tuan Dam, Kishan Panaganti, Brahim Driss et al.</dc:creator>
      <category>图像生成</category>
      <description>Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decision-making problems, yet it often relies on the assumption that the simulator and the real-world dynamics are identical....</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Online Robust Reinforcement Learning Through Monte-Carlo Planning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tuan Dam, Kishan Panaganti, Brahim Driss et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decision-making problems, yet it often relies on the assumption that the simulator and the real-world dynamics are identical. Although this assumption helps achieve the success of MCTS in games like Chess, Go, and Shogi, the real-world scenarios incur ambiguity due to their modeling mismatches in low-fidelity simulators. In this work, we present a new robust variant of MCTS that mitigates dynamical model ambiguities. Our algorithm addresses transition dynamics and reward distribution ambiguities to bridge the gap between simulation-based planning and real-world deployment. We incorporate a robust power mean backup operator and carefully designed exploration bonuses to ensure finite-sample convergence at every node in the search tree. We show that our algorithm achieves a convergence rate of $\mathcal{O}(n^{-1/2})$ for the value estimation at the root node, comparable to that of standard MCTS. Finally, we provide empirical evidence that our method achieves robust performance in planning problems even under significant ambiguity in the underlying reward distribution and transition dynamics.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18599v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18599v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] [图像生成] 4D Radar Perception Algorithms for Autonomous Driving: A Review</title>
      <link>https://arxiv.org/abs/2609.19216v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19216v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:46:07 GMT</pubDate>
      <dc:creator>Xumin Wu, Jun Zhou, Jilin Mei et al.</dc:creator>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Research on 4D millimeter-wave radar perception algorithms has flourished in recent years, extending from signal processing and object detection to semantic segmentation, motion estimation, occupancy ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">4D Radar Perception Algorithms for Autonomous Driving: A Review</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xumin Wu, Jun Zhou, Jilin Mei et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Research on 4D millimeter-wave radar perception algorithms has flourished in recent years, extending from signal processing and object detection to semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. This review organizes the field according to the evolution of perception tasks and algorithms. It first introduces radar fundamentals, data representations, and quality-enhancement methods, and then reviews object-level perception, motion and localization, local and dense spatial perception, and dynamic scene understanding. Across these directions, we compare radar-only learning, multimodal fusion, and cross-modal supervision and knowledge distillation. Particular attention is paid to how elevation, Doppler measurements, and radar physical priors are exploited across tasks. We further summarize the task coverage, input data, annotations, and evaluation protocols of existing datasets, clarifying the empirical support for different research directions. Finally, we discuss the common challenges and future directions of 4D radar perception for autonomous driving. This review provides a task-oriented perspective on the transition from sparse object perception to dynamic spatial understanding.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19216v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19216v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Peak-Aware Short-Term Load Forecasting Across Distribution Grid Aggregation Levels</title>
      <link>https://arxiv.org/abs/2609.18588v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18588v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:45:33 GMT</pubDate>
      <dc:creator>Souhardya Chattopadhyay, Julian Oelhaf, Antonia Schoening et al.</dc:creator>
      <category>模型架构</category>
      <description>For distribution system operators, short-term load forecasting (STLF) supports congestion management, voltage control, and asset protection. Most existing approaches focus on overall accuracy across a...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Peak-Aware Short-Term Load Forecasting Across Distribution Grid Aggregation Levels</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Souhardya Chattopadhyay, Julian Oelhaf, Antonia Schoening et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> mae</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">For distribution system operators, short-term load forecasting (STLF) supports congestion management, voltage control, and asset protection. Most existing approaches focus on overall accuracy across all time steps and neglect performance during high-demand (HD) periods, where larger forecast errors can increase the risk of congestion and voltage violations. In this paper, we study peak-aware STLF across three operator-relevant distribution grid aggregation levels, area codes (AC), secondary substations (SUB), and low-voltage (LV) feeders, using open datasets from the United Kingdom and Switzerland. We compare statistical baselines, machine learning models (LightGBM and XGBoost), and recent time-series foundation models (Chronos Bolt and Chronos-2) under a peak-aware evaluation framework that reports both overall and HD forecasting performance using NMAE and MAPE. The results show that Chronos-2 achieves the best HD performance across all aggregation levels, with HD-NMAE and HD-MAPE of 0.039 and 4.53% at AC, 0.080 and 9.45% at SUB, and 0.138 and 16.14% at LV, while Chronos-Bolt consistently ranks second best. Compared with the gradient boosted ML models, Chronos-2 reduces mean HD-NMAE by about 20-51% across levels while remaining best or near-best on the overall metrics. A quantile analysis of the probabilistic Chronos outputs further identifies aggregation-specific operating points, and runtime measurements indicate that foundation model inference is fast enough for practical deployment. Overall, the findings highlight peak-aware evaluation and aggregation specific quantile selection as a practical pathway toward more operationally relevant STLF in distribution networks.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18588v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18588v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking</title>
      <link>https://arxiv.org/abs/2609.18585v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18585v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:44:11 GMT</pubDate>
      <dc:creator>Giorgia Adorni, Michela Papandrea, Battista Rimoldi et al.</dc:creator>
      <category>模型架构</category>
      <description>Text-to-music (TTM) systems are increasingly used to generate musical audio from natural-language descriptions. Robust evaluation is therefore essential, yet reliable performance comparison remains ch...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Giorgia Adorni, Michela Papandrea, Battista Rimoldi et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Text-to-music (TTM) systems are increasingly used to generate musical audio from natural-language descriptions. Robust evaluation is therefore essential, yet reliable performance comparison remains challenging. This difficulty stems from differences in system architecture, supported conditioning information, and access mode, as well as heterogeneous and fragmented metrics that cannot be applied uniformly across systems. To address these challenges, we introduce TTM-Bench, a framework that defines a common protocol for systematic, reproducible performance benchmarking of contemporary TTM systems. It evaluates performance along two dimensions: musical-content alignment, quantified by interpretable semantic, genre, and musical-descriptor agreement scores against a common musical specification and summarized by an aggregate score; and computational efficiency, characterized by generation latency and real-time factor, alongside resource use for local models and cost for hosted services. We demonstrate the framework through a preliminary comparative case study, illustrating the complementary evidence captured by these dimensions. The results show that higher musical-content alignment does not systematically coincide with lower computational demands, highlighting the importance of assessing TTM performance through distinct, interpretable measures rather than a reductive overall indicator.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18585v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18585v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] On-the-Fly Homographies Calibration for Multi-Camera Tracking</title>
      <link>https://arxiv.org/abs/2609.18582v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18582v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:41:55 GMT</pubDate>
      <dc:creator>David Voihanski, Mor Sinai, Ben Zion Bobrovsky</dc:creator>
      <category>模型架构</category>
      <description>Precise multi-camera tracking traditionally relies on rigorous 3D site calibration, yet this requirement is often operationally impossible in large-scale deployments. Privacy regulations frequently pr...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">On-the-Fly Homographies Calibration for Multi-Camera Tracking</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> David Voihanski, Mor Sinai, Ben Zion Bobrovsky</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Precise multi-camera tracking traditionally relies on rigorous 3D site calibration, yet this requirement is often operationally impossible in large-scale deployments. Privacy regulations frequently prohibit recording video for offline calibration; limited bandwidth precludes synchronizing high-resolution streams from hundreds of cameras; and covering immense physical sites with calibration targets is logistically infeasible. We present a multi-camera homography calibration system designed to overcome these barriers through &quot;on-the-fly&quot; geometric refinement. Starting from coarse manual homographies, we introduce a centroid-based projection optimization (PO) that continuously aligns the ground-plane geometry using live detection streams. Because PO operates asynchronously on already-transmitted, lightweight metadata, it adds zero computational latency to the real-time tracker. This allows the system to adapt automatically to camera movements or environmental changes without human intervention. This optimized geometry feeds a multi-camera bird&apos;s-eye-view (BEV) tracker that fuses detections and unifies trajectories across zones. Crucially, by operating strictly on live anonymous metadata, our solution ensures a privacy-safe, zero-overhead, and resilient tracking pipeline that maintains global consistency in dynamic environments where static, recorded-video calibration is impossible.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18582v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18582v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Learning Where to Focus: Self-Supervised Multi-Scale ViTs for Histopathology</title>
      <link>https://arxiv.org/abs/2609.18578v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18578v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:38:41 GMT</pubDate>
      <dc:creator>Anabel Stammer, Valay Bundele, Mehran Hosseinzadeh et al.</dc:creator>
      <category>模型架构</category>
      <description>Pathologists diagnose diseases by first locating suspicious tissue and then examining it at higher magnification, whereas self-supervised vision transformers (ViTs) allocate the same spatial resolutio...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Learning Where to Focus: Self-Supervised Multi-Scale ViTs for Histopathology</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Anabel Stammer, Valay Bundele, Mehran Hosseinzadeh et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Pathologists diagnose diseases by first locating suspicious tissue and then examining it at higher magnification, whereas self-supervised vision transformers (ViTs) allocate the same spatial resolution to every image region despite diagnostic evidence being sparse and spanning multiple biological scales. Recent pathology foundation models have substantially improved representation quality by scaling training data and model capacity, but largely retain uniform tokenization. We instead investigate whether pathology representations can be improved by learning where to allocate spatial resolution during self-supervised learning. To this end, we propose CRAFT (Coarse-to-fine Region-Adaptive Feature Tokenization), a DINO-based framework that learns image-dependent mixed-scale representations by using self-supervised attention to selectively refine informative regions while preserving coarse context, together with a symmetric cross-scale regularization objective that encourages complementary coarse and fine representations. Across CAMELYON16, TCGA-Lung subtype classification, and TCGA-LUAD survival prediction, CRAFT consistently outperforms comparable-scale self-supervised methods while requiring lower inference computation. Despite using only a compact 22M parameter backbone trained on comparatively small pathology datasets, CRAFT remains competitive with, and often surpasses, substantially larger pathology foundation models.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18578v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18578v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Perceptual Refinement of an End-to-End Video Streaming Pipeline via Generative AI Layers</title>
      <link>https://arxiv.org/abs/2609.19215v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.19215v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:33:08 GMT</pubDate>
      <dc:creator>Emanuele Artioli, Farzad Tashtarian, Christian Timmerer</dc:creator>
      <category>模型架构</category>
      <description>Traditional codecs treat every region of a frame alike; a generative layer can instead degrade the regions a viewer attends to least and reconstruct them at the client. We present PRESLEY, which exten...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Perceptual Refinement of an End-to-End Video Streaming Pipeline via Generative AI Layers</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Emanuele Artioli, Farzad Tashtarian, Christian Timmerer</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Traditional codecs treat every region of a frame alike; a generative layer can instead degrade the regions a viewer attends to least and reconstruct them at the client. We present PRESLEY, which extends the prior conference work ELVIS by replacing destructive block removal with adaptive in-place degradation under a removability mask, signaling per-block strength in a bit-packed side channel, and restoring via generative backbones conditioned on transmitted visual priors rather than unconditioned in-painting. We separate the problem into three goals: choosing which blocks to degrade, degrading them so the encoder spends fewer bits, and restoring them. Against its predecessor at matched rate, PRESLEY achieves a decisive mean -56.4% BD-rate reduction on delivered background quality across 13 rate ladders spanning multiple codecs and dataset families. Against pristine baselines, PRESLEY defines the operating regime of generative transport: delivering substantial bitrate savings (up to -29.4% BD-rate) and superior background quality (17/23 sequences) in the target bit-starved regime, while maintaining foreground fidelity bit-exact. We further map where the theoretical headroom in this class of architecture lies. Using an exact leave-one-superblock-out combinatorial oracle as an additive empirical bound, we show that existing complexity heuristics already capture 83.3% of bit-cost savings, bounding remaining cost-axis headroom at about 5% of total bitrate. We then identify and model the primary unaddressed axis -- post-restoration damage -- which disperses widely (4.9-8.4 dB). We prove that this damage is predictable before transmission (held-out rho = +0.400), establishing the feasibility of transmit-time restorability modeling and defining the roadmap for joint rate-distortion-restoration selection rules.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.19215v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.19215v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] CARA: Collision-Aware Resolution Adaptation for Multiresolution Hash Encoding Based Image Fitting</title>
      <link>https://arxiv.org/abs/2609.18554v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18554v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:17:34 GMT</pubDate>
      <dc:creator>Linfeng Ye, Zhixiang Chi, Shayan Mohajer Hamidi et al.</dc:creator>
      <category>模型架构</category>
      <description>Multiresolution hash encodings have recently enabled fast and high-fidelity implicit neural representations by storing multi-scale features in fixed-size hash tables along a geometric resolution sched...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">CARA: Collision-Aware Resolution Adaptation for Multiresolution Hash Encoding Based Image Fitting</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Linfeng Ye, Zhixiang Chi, Shayan Mohajer Hamidi et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Multiresolution hash encodings have recently enabled fast and high-fidelity implicit neural representations by storing multi-scale features in fixed-size hash tables along a geometric resolution schedule. However, the standard design is data-agnostic: different resolution levels receive identical hash-table capacity despite large differences in image frequency content. As a result, some levels experience severe hash collisions while others underutilize parameters, leading to inefficient capacity allocation. To address this issue, we propose Collision-Aware Resolution Adaptation (CARA), a method that assigns per-level resolutions by balancing the effective information load across hash levels. This adaptive allocation reduces capacity bottlenecks and improves parameter efficiency. In addition, we introduce an invertible pixel-shuffle transform that reduces hash load factors by redistributing spatial information, thereby mitigating collision-induced information loss without enlarging the hash tables. To support evaluation on extremely high-resolution data, we also curate, to the best of our knowledge, the first uncompressed whole-slide image dataset for academic research. Experiments on Kodak images, gigapixel natural images, and raw whole-slide images demonstrate that CARA consistently improves the fidelity-parameter trade-off. Our method matches state-of-the-art performance while using only $27.76%$ of the parameters, and achieves up to $6.11$ dB PSNR improvement at comparable parameter counts. Code is provided in the supplementary.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18554v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18554v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] HAP: A Hand-Driven Active Perception Framework for Egocentric Head Motion Prediction</title>
      <link>https://arxiv.org/abs/2609.18548v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18548v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:13:28 GMT</pubDate>
      <dc:creator>Yunji Feng, Junyi Ma, Guanzhong Sun et al.</dc:creator>
      <category>模型架构</category>
      <description>Egocentric motion forecasting has primarily focused on hands and manipulated objects, leaving future human head motion comparatively underexplored. During manipulation, the head both redirects percept...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">HAP: A Hand-Driven Active Perception Framework for Egocentric Head Motion Prediction</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yunji Feng, Junyi Ma, Guanzhong Sun et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Egocentric motion forecasting has primarily focused on hands and manipulated objects, leaving future human head motion comparatively underexplored. During manipulation, the head both redirects perception toward the target to acquire task-relevant evidence and coordinates with body and hand motion. We therefore formulate future six Degree of Freedom (6-DoF) head-motion prediction conditioned on observed hand motion and inferred target context, and propose HAP, a Hand-Driven Active Perception framework. HAP infers confidence for each target object from observed hand motion and object geometry. Then constructs a dynamic Predictive Target-Centric Amodal Occlusion Graph (P-TAOG) representing current and potential occlusion among candidate objects. Directed graph and causal temporal reasoning encode the evolving target conditioned perceptual state, which is fused with hand and head motion history. A horizon-wise gate then blends the learned trajectory with a constant velocity prior. We further introduce Bottle, an egocentric RGB-D dataset of object manipulation toward specified targets, with coordinated head and hand motion under changing target visibility. Experiments on the public dataset and Bottle show that HAP achieves lower head motion prediction errors than representative baselines, supporting the value of hand driven intention and dynamic occlusion reasoning for anticipating human head motion. Code will be released at https://HAP-ego.github.io/HAP.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18548v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18548v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] STUNet-Fusion: Spatiotemporal Needle-Tip Localization in Ultrasound Video via Multi-Channel Motion Fusion</title>
      <link>https://arxiv.org/abs/2609.18546v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18546v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:11:18 GMT</pubDate>
      <dc:creator>Chia-Chi Hsu, Chia-Hsuan Hsu, Che-Chou Shen</dc:creator>
      <category>模型架构</category>
      <description>Needle-tip localization in ultrasound remains challenging because the needle may appear weak, discontinuous, or partially invisible, while imaging artifacts and anatomical structures can produce simil...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">STUNet-Fusion: Spatiotemporal Needle-Tip Localization in Ultrasound Video via Multi-Channel Motion Fusion</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Chia-Chi Hsu, Chia-Hsuan Hsu, Che-Chou Shen</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> u-net</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Needle-tip localization in ultrasound remains challenging because the needle may appear weak, discontinuous, or partially invisible, while imaging artifacts and anatomical structures can produce similar responses. To address this problem, we propose STUNet-Fusion, a spatiotemporal framework for needle-tip localization in ultrasound videos. The proposed method formulates the input as a tri-channel spatio-temporal fusion tensor, comprising grayscale appearance, grid-based motion feature, and raw frame difference. A shared ResNet-34 encoder extracts spatial features, ConvLSTM integrates temporal dependencies, and a U-Net decoder reconstructs a dense probability heatmap. The final coordinates are extracted via a soft-argmax operation to achieve sub-pixel localization accuracy. Experimental results demonstrate that this spatiotemporal fusion strategy significantly improves localization robustness compared to conventional baselines.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18546v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18546v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Accuracy- and Real-Time-Aware 4D Radar Preprocessing for Autonomous Driving Perception Systems</title>
      <link>https://arxiv.org/abs/2609.18542v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18542v1</guid>
      <pubDate>Wed, 16 Sep 2026 12:06:22 GMT</pubDate>
      <dc:creator>Woo-Jin Jung, Dong-Hee Paek, Jeong-Su Park et al.</dc:creator>
      <category>模型架构</category>
      <description>4D radar has emerged as a promising next-generation sensor for improving the robustness of autonomous driving perception systems because of its stable sensing capability under adverse weather conditio...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Accuracy- and Real-Time-Aware 4D Radar Preprocessing for Autonomous Driving Perception Systems</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Woo-Jin Jung, Dong-Hee Paek, Jeong-Su Park et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">4D radar has emerged as a promising next-generation sensor for improving the robustness of autonomous driving perception systems because of its stable sensing capability under adverse weather conditions. However, deploying 4D radar in embedded environments with limited hardware resources requires radar-representation preprocessing that jointly considers perception accuracy, real-time performance, and computational complexity. This paper proposes a preprocessing framework for 4D-radar-based 3D object detection. First, Percentile-based 3D Shape Preservation (P3DP) extracts point clouds from radar tensors while preserving object-shape information and suppressing noise and false alarms. Second, Multi-frame-based Noise Point Discrimination using Kernel Density Estimation (MF-KDE) improves the density and reliability of sparse radar point clouds. Finally, Embedded \&amp; NetScore (ENS) evaluates suitability for embedded deployment by jointly considering accuracy, real-time performance, adverse-weather robustness, and model complexity.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18542v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18542v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] DiT-Garment: Garment Dynamics with Diffusion Transformers</title>
      <link>https://arxiv.org/abs/2609.18510v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18510v1</guid>
      <pubDate>Wed, 16 Sep 2026 11:44:03 GMT</pubDate>
      <dc:creator>Antoine Dumoulin, Laurence Boissieux, Joao Regateiro et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>We present DiT-Garment to model dynamic 3D clothing over human body models in arbitrary motion. Unlike existing methods, DiT-Garment can animate garments with unseen designs and physical materials, wh...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">DiT-Garment: Garment Dynamics with Diffusion Transformers</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Antoine Dumoulin, Laurence Boissieux, Joao Regateiro et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We present DiT-Garment to model dynamic 3D clothing over human body models in arbitrary motion. Unlike existing methods, DiT-Garment can animate garments with unseen designs and physical materials, while allowing for direct inference of deformations for any target pose. To achieve this, we leverage a 2D diffusion transformer architecture to learn 3D deformations in a 2D UV-space. As the result is non-deterministic, our generative model learns the distribution of possible outcomes. The template garment is represented as a 3D triangle mesh spatially aligned with a 3D human body model in a standardized pose. To work with different garment designs without the need of a common template or complex graph convolution operations, the diffusion transformer is conditioned on a 3D position map of the template, represented in UV-space, which allows to implicitly learn a deformation of the 3D space around the body in standard pose. Further conditioning on body motion and physical parameters allows to physically ground the model. We quantitatively and qualitatively evaluate DiT-Garment on both synthetic and real data. While only trained on synthetic simulations of automatically generated cloth designs, our method generalizes to captured and artist-made garment designs. Code and data are available for research purposes at https://dumoulina.github.io/dit-garment/.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18510v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18510v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] [图像生成] Learning A Unified Template for Gait Recognition</title>
      <link>https://arxiv.org/abs/2609.18490v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18490v1</guid>
      <pubDate>Wed, 16 Sep 2026 11:27:36 GMT</pubDate>
      <dc:creator>Panjian Huang, Saihui Hou, Junzhou Huang et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>&quot;What I cannot create, I do not understand.&quot;Human wisdom reveals that creation is one of the highest forms of learning. For example, Diffusion Models have demonstrated remarkable semantic structure an...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Learning A Unified Template for Gait Recognition</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Panjian Huang, Saihui Hou, Junzhou Huang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, diffusion model, image generation, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">&quot;What I cannot create, I do not understand.&quot;Human wisdom reveals that creation is one of the highest forms of learning. For example, Diffusion Models have demonstrated remarkable semantic structure and memory in image generation, understanding, and restoration, which intuitively benefits representation learning. However, current gait networks rarely embrace this perspective, relying primarily on learning by contrasting gait samples under varying complex conditions, leading to semantic inconsistency and uniformity issues. To address these issues, we propose Origins with generative capabilities whose underlying philosophy is that different entities are generated from a unified template, inherently regularizing gait representations within a consistent and diverse semantic space to capture accurate gait differences. Admittedly, learning this unified template is exceedingly challenging, as it requires the comprehensiveness of the template to encompass gait representations with various conditions. Inspired by Diffusion Models, Origins diffuses the unified template into timestep templates for gait generative learning, and meanwhile transfers the unified template for gait representation learning. Especially, gait generative and representation learning serve as a unified framework for end-to-end joint training. Extensive experiments on CASIA-B, CCPG,SUSTech1K, Gait3D, GREW and CCGR-MINI demonstrate that Origins performs unified generative and representation learning, achieving superior performance.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18490v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18490v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] Beyond Random Couplings: Contrastive Noise Alignment in Generative Flows</title>
      <link>https://arxiv.org/abs/2609.18488v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18488v1</guid>
      <pubDate>Wed, 16 Sep 2026 11:21:31 GMT</pubDate>
      <dc:creator>Lennart Wittke, Vinicius Azevedo</dc:creator>
      <category>扩散模型</category>
      <description>Diffusion and flow-matching models are typically trained by corrupting data through independently sampled Gaussian noise. While simple and scalable, this forward process induces arbitrary data-noise c...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Beyond Random Couplings: Contrastive Noise Alignment in Generative Flows</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Lennart Wittke, Vinicius Azevedo</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, rectified flow</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Diffusion and flow-matching models are typically trained by corrupting data through independently sampled Gaussian noise. While simple and scalable, this forward process induces arbitrary data-noise couplings, forcing the network to learn high-curvature transports between unrelated endpoints. Existing optimal-transport methods reduce this burden by reassigning fixed noise samples to data, but the source noise distribution itself remains passive. To address this, we introduce Contrastive Noise Alignment (CNA), a training-time method that creates dynamic, contrastive couplings by optimizing the noise representations directly. By modeling the noise batch as an interacting particle system, CNA employs a cross-modal InfoNCE objective to align noise particles with their paired data targets. To prevent spatial collapse, this alignment is regularized using an angular entropy term and a radial norm penalty. We show theoretically that this equilibrium asymptotically preserves Gaussian structures, maintaining tractability during inference. Empirically, CNA improves the alignment between noise and data, reduces flow curvature, and provides better generation quality with fewer required sampling steps. For few-step, pixel-space generation (2-4 NFEs), CNA reduces FID by over 50\% compared to standard rectified flow, and by at least 24\% against Optimal Transport baselines.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18488v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18488v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models</title>
      <link>https://arxiv.org/abs/2609.18487v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18487v1</guid>
      <pubDate>Wed, 16 Sep 2026 11:19:54 GMT</pubDate>
      <dc:creator>Shijie Lian, Bin Yu, Zhaolong Shen et al.</dc:creator>
      <category>模型架构</category>
      <description>Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted token...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shijie Lian, Bin Yu, Zhaolong Shen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18487v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18487v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] GeoCond: A Conditioning-Aware Reliability Adapter for Feed-Forward 3D Reconstruction</title>
      <link>https://arxiv.org/abs/2609.18465v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18465v1</guid>
      <pubDate>Wed, 16 Sep 2026 11:00:01 GMT</pubDate>
      <dc:creator>David Ahmedt-Aristizabal, Mohammad Ali Armin, Russell Tsuchida et al.</dc:creator>
      <category>模型架构</category>
      <description>Feed-forward 3D foundation models such as VGGT predict cameras, depth, and point maps in a single pass, but can fail silently under low overlap, low parallax, and extreme relative rotation. Stratified...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">GeoCond: A Conditioning-Aware Reliability Adapter for Feed-Forward 3D Reconstruction</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> David Ahmedt-Aristizabal, Mohammad Ali Armin, Russell Tsuchida et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Feed-forward 3D foundation models such as VGGT predict cameras, depth, and point maps in a single pass, but can fail silently under low overlap, low parallax, and extreme relative rotation. Stratified analyses over these factors show that these failures are governed by geometric conditioning and are poorly captured by native aleatoric confidence. We introduce GeoCond, a lightweight reliability adapter for frozen feed-forward 3D backbones. GeoCond reads the backbone&apos;s predicted geometry and outputs pose-level uncertainty and a refinement gate. During training, it can be supervised by frame-permutation orbit variance, ground-truth pose error when labels are available, or cycle residuals from unlabelled independent pose graphs. At inference, the default head requires only one backbone pass and a small MLP. On VGGT, GeoCond improves out-of-distribution (OOD) AUSE (area under the sparsification-error curve; lower is better) from $0.32$ to $0.20$ over native confidence, transfers zero-shot to outdoor extreme-view scenes, and avoids the collapse caused by applying bundle adjustment uniformly. Across multiple backbones, cycle-distilled variants provide a ground-truth-free adaptation route, including cases where permutation variance vanishes on equivariant models. The same reliability signal supports gated refinement, pose-graph weighting, calibration, curation, and capture decisions. Reliable feed-forward 3D reconstruction requires not only predicting geometry, but also knowing when that geometry should be trusted.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18465v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18465v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models</title>
      <link>https://arxiv.org/abs/2609.18462v2</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18462v2</guid>
      <pubDate>Wed, 16 Sep 2026 10:56:39 GMT</pubDate>
      <dc:creator>Tianbin Liu, Jian Zhu, Taiyi Su et al.</dc:creator>
      <category>模型架构</category>
      <description>FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-spe...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tianbin Liu, Jian Zhu, Taiyi Su et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18462v2" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18462v2" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] [图像生成] DR.WILSS: Diffusion-Based Replay for Weakly Supervised Continual Semantic Segmentation</title>
      <link>https://arxiv.org/abs/2609.18444v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18444v1</guid>
      <pubDate>Wed, 16 Sep 2026 10:33:51 GMT</pubDate>
      <dc:creator>Leon Arthur Marx, Francesco Barbato, Matteo Caligiuri et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Weakly supervised class-incremental semantic segmentation (WILSS) aims to train a segmentation model over multiple steps, each introducing new concepts to be learned with only image-level supervision....</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">DR.WILSS: Diffusion-Based Replay for Weakly Supervised Continual Semantic Segmentation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Leon Arthur Marx, Francesco Barbato, Matteo Caligiuri et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, lora, dit, inpainting</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Weakly supervised class-incremental semantic segmentation (WILSS) aims to train a segmentation model over multiple steps, each introducing new concepts to be learned with only image-level supervision. We introduce DR$.$WILSS, an innovative approach to address catastrophic forgetting in continual learning using diffusion-based generative replay. Our framework leverages language clues to guide the diffusion process, employing self-inpainting and regularization techniques to efficiently produce replay data, aiding the learning process. By generating high-quality replay data, the information from previously learned classes can be preserved during continual updates, a critical challenge in incremental learning scenarios. To further align the statistics of replay data with those of training samples, we apply LoRAs to the generative model. Experimental results demonstrate state-of-the-art performance across multiple benchmarks and generative architectures, while avoiding storage of training data and the use of additional resource-demanding tools during training. The proposed technique enables an optimal tradeoff between training complexity and inference-time accuracy, making DR$.$WILSS a promising solution for real-world applications.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18444v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18444v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective</title>
      <link>https://arxiv.org/abs/2609.18432v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18432v1</guid>
      <pubDate>Wed, 16 Sep 2026 10:24:31 GMT</pubDate>
      <dc:creator>Panjian Huang, Yunjie Peng, Saihui Hou et al.</dc:creator>
      <category>模型架构</category>
      <description>Extensive occlusions in real-world scenarios pose challenges to gait recognition due to missing and noisy information, as well as body misalignment in position and scale. We argue that rich dynamic co...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Panjian Huang, Yunjie Peng, Saihui Hou et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> mae, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Extensive occlusions in real-world scenarios pose challenges to gait recognition due to missing and noisy information, as well as body misalignment in position and scale. We argue that rich dynamic contextual information within a gait sequence inherently possesses occlusion-solving traits: 1) Adjacent frames with gait continuity allow holistic body regions to infer occluded body regions; 2) Gait cycles allow information integration between holistic actions and occluded actions. Therefore, we introduce an action detection perspective where a gait sequence is regarded as a composition of actions. To detect accurate actions under complex occlusion scenarios, we propose an Action Detection Based Mixture of Experts (GaitMoE), consisting of Mixture of Temporal Experts (MTE) and Mixture of Action Experts (MAE). MTE adaptively constructs action anchors by temporal experts and MAE adaptively constructs action proposals from action anchors by action experts. Especially, action detection as a proxy task with gait recognition is an end-to-end joint training only with ID labels. In addition, due to the lack of a unified occluded benchmark, we construct a pioneering Occluded Gait database (OccGait), containing rich occlusion scenarios and annotations of occlusion types. Extensive experiments on OccGait, OccCASIA-B,Gait3D and GREW demonstrate the superior performance of GaitMoE.OccGait is available at https://github.com/BNU-IVC/OccGait.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18432v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18432v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [模型架构] [扩散模型] StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions</title>
      <link>https://arxiv.org/abs/2609.18430v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18430v1</guid>
      <pubDate>Wed, 16 Sep 2026 10:23:00 GMT</pubDate>
      <dc:creator>Awomo-WM Team, :, Enhui Ma et al.</dc:creator>
      <category>视频生成</category>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucPhysVideo, a family of video world models that bri...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Awomo-WM Team, :, Enhui Ma et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, video prediction, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucPhysVideo, a family of video world models that bridges physics-focused data curation with language- and action-conditioned prediction of scene evolution. Our data pipeline combines motion-aware video segmentation, quality and content filtering, and physical relevance verification with structured annotations of objects, materials, and temporally localized interactions. By disentangling camera motion from object behavior and explicitly describing contact, deformation, and state transitions, the pipeline provides supervision grounded in observable physical events. Building on these data, we introduce StrucPhysVideo-TI2V, a sparse Mixture-of-Experts (MoE) text-image-to-video model trained with a curriculum that progressively emphasizes physical dynamics while retaining general-domain video data. StrucPhysVideo-TI2V achieves state-of-the-art performance on Physics-IQ Verified, scoring 45.5% and outperforming Cosmos3-Super-Image2Video by 2.8 percentage points. Caption ablations across backbones further demonstrate the effectiveness of physics-focused supervision. We further extend StrucPhysVideo-TI2V to StrucPhysVideo-IA2V, an interactive image-action-to-video world model that predicts visual outcomes from robot end-effector commands. Action conditioning, causal autoregressive generation, and few-step distillation enable incremental robot rollouts with only four denoising steps. Together, StrucPhysVideo advances physical dynamics modeling from image- and language-conditioned video prediction toward action-driven interaction.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18430v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18430v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] Vocabulary-Guided Gait Recognition</title>
      <link>https://arxiv.org/abs/2609.18413v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18413v1</guid>
      <pubDate>Wed, 16 Sep 2026 10:07:20 GMT</pubDate>
      <dc:creator>Panjian Huang, Saihui Hou, Chunshui Cao et al.</dc:creator>
      <category>多模态生成</category>
      <description>What is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a gait as the human inherent structure from points...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Vocabulary-Guided Gait Recognition</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Panjian Huang, Saihui Hou, Chunshui Cao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">What is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a gait as the human inherent structure from points. However, the considerations remain vague for humans to comprehend truly. In this work, we introduce a novel paradigm Vocabulary-Guided Gait Recognition, dubbed Gait-World, which attempts to explore gait concepts through human vocabularies with Vision-Language Models (VLMs). Although VLMs have achieved the remarkable progress in various vision tasks, the cognitive capability regarding gait modalities remains limited. The success element in Gait-World is the proper vocabulary prompt where this paradigm carefully selects gait cycle actions as Vocabulary Base, bridging the gait and vocabulary feature spaces and further promoting human understanding for the gait. How to extract gait features? Although previous gait networks have made significant progress, learning solely from gait modalities on limited gait databases makes it difficult to learn universal gait features for practicality. Therefore, we propose the first Gait-World model, dubbed α-Gait, which guides the gait network learning with vocabulary knowledge from VLMs. However, due to the heterogeneity of the modalities, directly integrating vocabulary and gait features is highly challenging as they reside in different embedding spaces. To address the issues, α-Gait designs Vocabulary Relation Mapper and Gait Fine grained Detector to map and establish vocabulary relations in the gait space for detecting corresponding gait features. Extensive experiments on CASIA-B, CCPG, SUSTech1K, Gait3D and GREW reveal the potential value and research directions of vocabulary information from VLMs in the gait field.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18413v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18413v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Prosthesis-Aware 3D Human Pose Estimation: A Dataset and Benchmark for RSP Users</title>
      <link>https://arxiv.org/abs/2609.18406v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18406v1</guid>
      <pubDate>Wed, 16 Sep 2026 09:59:54 GMT</pubDate>
      <dc:creator>Yilin Wen, Kechuan Dong, Fumiya Suginaka et al.</dc:creator>
      <category>模型架构</category>
      <description>Recovering 3D human body motion from video is important for applications such as rehabilitation assessment and sports performance evaluation. For prosthesis users, this requires capturing both natural...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Prosthesis-Aware 3D Human Pose Estimation: A Dataset and Benchmark for RSP Users</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yilin Wen, Kechuan Dong, Fumiya Suginaka et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Recovering 3D human body motion from video is important for applications such as rehabilitation assessment and sports performance evaluation. For prosthesis users, this requires capturing both natural body joints and the geometry of the prosthetic device, a challenge that existing methods are not designed to address. Model-based estimators rely on body models trained on non-amputee individuals and cannot represent prosthesis geometry, while model-free methods lack body kinematic priors and are unreliable under occlusion. This challenge is particularly prominent for users of running-specific prostheses (RSPs), where the RSP has a complex curved geometry and moves dynamically during exercise. To fill this gap, we collect RSP3D, the first 3D dataset of RSP users, covering essential daily-life and exercise actions from participants with varied amputation conditions, using a multi-camera marker-based motion capture setup. We formally define the task of prosthesis-aware 3D pose estimation, evaluate representative methods in a zero-shot setting, and confirm their individual limitations. We further propose a hybrid baseline combining model-based body joint estimation with model-free RSP shape recovery, establishing a starting point for future research.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18406v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18406v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] A Non-Linear Neuron Based Detection of Isolated Pixels in Binary and Grayscale Images using Contrast Sensitive Receptive Fields</title>
      <link>https://arxiv.org/abs/2609.18399v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18399v1</guid>
      <pubDate>Wed, 16 Sep 2026 09:56:03 GMT</pubDate>
      <dc:creator>Nassir Mohammad</dc:creator>
      <category>图像生成</category>
      <description>Identifying isolated points is important in image processing applications such as medical imaging, astronomy and quality control management. Other domains, such as cybersecurity, also present challeng...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Non-Linear Neuron Based Detection of Isolated Pixels in Binary and Grayscale Images using Contrast Sensitive Receptive Fields</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Nassir Mohammad</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Identifying isolated points is important in image processing applications such as medical imaging, astronomy and quality control management. Other domains, such as cybersecurity, also present challenges that can be framed as image processing problems. One example of particular interest is the identification of anomalous single nodes in spatially organised networks where groups of nodes in different regions share similar feature values. This task can involve both binary and more complex grayscale images. However, existing methods face limitations: template matching is infeasible for grayscale images, while 2nd order derivative based methods are highly sensitive to noise and require user-specified thresholds. To overcome these issues, a novel method is proposed for detecting meaningful single-pixel deviations in images. This approach modifies and extends a neuron model, originally designed for anomaly detection, to operate on spatially diameter limited receptive fields that incorporate excitatory and inhibitory regions. The result is a method that is free from user-specified thresholds and parameters, and can be applied to both binary and grayscale images, providing an effective, robust and efficient solution.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18399v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18399v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [模型架构] MSR: Multiple Subject Reference for Video Generation</title>
      <link>https://arxiv.org/abs/2609.18393v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18393v1</guid>
      <pubDate>Wed, 16 Sep 2026 09:51:50 GMT</pubDate>
      <dc:creator>Guannan Li, Jiaji Chen, Jingyuan Liao et al.</dc:creator>
      <category>视频生成</category>
      <category>模型架构</category>
      <description>Conditioning a video generator on multiple images requires preserving appearance while associating each reference with its intended role. We present MSR (Multiple Subject Reference), a slot-aware cond...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MSR: Multiple Subject Reference for Video Generation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Guannan Li, Jiaji Chen, Jingyuan Liao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> video generation, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Conditioning a video generator on multiple images requires preserving appearance while associating each reference with its intended role. We present MSR (Multiple Subject Reference), a slot-aware conditioning scheme for LTX-based video generation. Each reference image is independently encoded as a static clip and represented by a separate latent-token group. A compact Fourier-feature multilayer perceptron adds a numeric slot embedding, while slot-dependent temporal offsets modify the group&apos;s rotary coordinates. The reference groups are prepended to noisy target tokens and serve as clean context during target-only flow-matching training. We implement this scheme through low-rank adaptation and release the resulting weights and inference workflows. Qualitative examples demonstrate compositions containing distinct characters and referenced environments in realistic and stylized scenes. Development observations suggest reduced reference confusion relative to an earlier continuous-reference baseline, while similar clothing, complex garments, and viewpoint changes remain challenging. We describe the conditioning mechanism, the retained training configuration, and the observed strengths and limitations of the released system. A supplementary audio-reference experiment adds voice conditioning while keeping the visual parameters frozen.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18393v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18393v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] JigSync: Gauge-Resolved Synchronization for Jigsaw Reassembly under Unknown Piece Orientation</title>
      <link>https://arxiv.org/abs/2609.18379v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18379v1</guid>
      <pubDate>Wed, 16 Sep 2026 09:34:31 GMT</pubDate>
      <dc:creator>Soham Pahari, Antik Aich Roy, Ujjwal Bhattacharya</dc:creator>
      <category>模型架构</category>
      <description>Square jigsaw reassembly requires recovering the spatial arrangement of shuffled fragments from their visual content and pairwise relationships. While recent studies have made substantial progress, ex...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">JigSync: Gauge-Resolved Synchronization for Jigsaw Reassembly under Unknown Piece Orientation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Soham Pahari, Antik Aich Roy, Ujjwal Bhattacharya</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Square jigsaw reassembly requires recovering the spatial arrangement of shuffled fragments from their visual content and pairwise relationships. While recent studies have made substantial progress, existing benchmarks typically assume that all fragments are provided upright, reducing reassembly to a permutation problem. We study the generalized problem in which each fragment may also have gone through an unknown rotation. For this setting we establish a gauge-unobservability theorem: the minimum of the weighted least-squares objective is exactly invariant under a uniform global rotation of arbitrary magnitude, so no residual-based criterion can recover the global orientation. The theorem further identifies how the issue of global orientation can be resolved: an orientation anchor estimated from the content of a single fragment, lying outside its scope, suffices. To address the above, we propose JigSync, which attains 63.8% and 31.8% absolute accuracy (AA) on GAP-3 and GAP-5, respectively, the highest reported on both, while additionally recovering a rotation per piece that neither benchmark requires. We release JigSync, a degradation protocol that sweeps shape, erosion, photometry, grid size, and rotation independently.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18379v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18379v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] Visual Input and Its Framing Affect Attribute-based Descriptions Produced by Large Vision-Language Models</title>
      <link>https://arxiv.org/abs/2609.18345v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18345v1</guid>
      <pubDate>Wed, 16 Sep 2026 09:08:57 GMT</pubDate>
      <dc:creator>Xiaomeng Wang, Martha Larson, Zhengyu Zhao</dc:creator>
      <category>多模态生成</category>
      <description>Large vision-language models (LVLMs) are commonly used with only a single text prompt as the input, or plus an image. In this paper, we demonstrate that when the image exists, even if the text prompt ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Visual Input and Its Framing Affect Attribute-based Descriptions Produced by Large Vision-Language Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xiaomeng Wang, Martha Larson, Zhengyu Zhao</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Large vision-language models (LVLMs) are commonly used with only a single text prompt as the input, or plus an image. In this paper, we demonstrate that when the image exists, even if the text prompt is not about the specific instance (but only the concept it belongs to) in that image, the response would still be affected. For example, when the text prompt only asks for the attribute descriptions of a dog breed, an image depicting a specific dog from that breed would shift the response. Further, how the specific instance is framed in that image would determine towards which the response shifts. Detailed analyses also reveal that in the response, physical terms increase from 18% for text-only to 45% (40%) for subject-focused (subject-in-situation) framings. Overall, the unexpected effects of visual cues on LVLMs highlight the need to understand the presence of an image and its framing when evaluating the robustness of LVLMs.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18345v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18345v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Pose2Muscle: Structured Spatio-Temporal Decoding for Discrete Muscle Activity Estimation from Human Pose</title>
      <link>https://arxiv.org/abs/2609.18336v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18336v1</guid>
      <pubDate>Wed, 16 Sep 2026 08:58:45 GMT</pubDate>
      <dc:creator>Yuepeng Chen, Jiehong Shi, Kaili Zheng et al.</dc:creator>
      <category>模型架构</category>
      <description>Muscle activity is fundamental to human movement, and understanding its patterns is critical for injury prevention and rehabilitation. Conventional muscle activity monitoring relies on specialized sen...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Pose2Muscle: Structured Spatio-Temporal Decoding for Discrete Muscle Activity Estimation from Human Pose</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yuepeng Chen, Jiehong Shi, Kaili Zheng et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Muscle activity is fundamental to human movement, and understanding its patterns is critical for injury prevention and rehabilitation. Conventional muscle activity monitoring relies on specialized sensors such as surface electromyography, which limits its practicality for long-term real-world use. Existing studies suggest that muscle-related information can be inferred from human pose. However, the substantial gap between externally observable pose and internal muscle activation, limits the accuracy and generalization of current approaches. In this study, we propose Pose2Muscle, a pose-driven framework for discrete muscle activity estimation without requiring sEMG signals at inference time. Instead of directly regressing continuous sEMG signals, Pose2Muscle reformulates muscle estimation as a structured prediction problem over discrete muscle activity states, yielding a more stable and interpretable target space. The framework combines multi-scale spatio-temporal attention to capture motion patterns at complementary spatial and temporal scales with a directed acyclic graph-based decoder that maintains multiple candidate muscle-state hypotheses and performs structured trajectory inference over time. To support this task, we construct PoseEMG-43, a synchronized pose-sEMG dataset containing 2,992 movement instances from 43 daily-life actions performed by 14 participants. Experiments show that Pose2Muscle consistently outperforms representative retrieval- and pose-based baselines. It achieves an Adjacent-level Accuracy of 86.36% and a Pearson correlation coefficient of 0.8821 under the Random Split, and 63.97% and 0.6795, respectively, under the Subject-Level Split. These results demonstrate the feasibility of inferring structured muscle-state patterns from human pose and suggest the potential of Pose2Muscle for muscle-aware movement analysis when direct physiological sensing is impractical</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18336v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18336v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] PDA++: Field-Aligned Planning and Scene-Adaptive Insertion in Remote Sensing</title>
      <link>https://arxiv.org/abs/2609.18329v2</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18329v2</guid>
      <pubDate>Wed, 16 Sep 2026 08:53:30 GMT</pubDate>
      <dc:creator>Xianchi Dong, Yingyan Hou, Chao Ren et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Remote sensing recognition is often constrained by scarce observations of rare targets and costly annotations, making realistic synthetic augmentation particularly valuable for few-shot and long-taile...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PDA++: Field-Aligned Planning and Scene-Adaptive Insertion in Remote Sensing</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Xianchi Dong, Yingyan Hou, Chao Ren et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Remote sensing recognition is often constrained by scarce observations of rare targets and costly annotations, making realistic synthetic augmentation particularly valuable for few-shot and long-tailed scenarios. Object insertion provides an efficient way to increase target diversity while preserving authentic background scenes, but realistic insertion in overhead imagery requires the generated target to adapt coherently to its surrounding environment. To this end, we propose PDA++, a unified environment-aware object insertion framework organized as Plan, Decouple, and Assimilate. Planning determines scene-compatible poses through an affordance field that combines geometric clearance with structure- and scale-aware cues. Decoupling introduces a pose-conditioned background that provides precise spatial guidance together with target-scene context, allowing the reference object to preserve its identity while adapting to the target observation. This construction also naturally provides pixel-level masks for segmentation augmentation. Assimilation further improves local coherence by aligning multi-scale texture distributions through optimal transport. On the optical benchmark, PDA++ achieves a whole-image FID of 6.28 and improves average few-shot recognition mAP50 by 17.69 points, corresponding to a 28.8% relative gain over the real-data baseline. On SAR imagery, it improves ship detection by 4.10 mAP50 points and remains effective under cross-dataset transfer and amorphous-target insertion. Code is available at https://github.com/lisheyu972/PDA_PLUS.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18329v2" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18329v2" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [多模态生成] [图像生成] Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model</title>
      <link>https://arxiv.org/abs/2609.18323v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18323v1</guid>
      <pubDate>Wed, 16 Sep 2026 08:43:11 GMT</pubDate>
      <dc:creator>Haoyu Zhao, Zihao Zhao, Tianyu Deng et al.</dc:creator>
      <category>视频生成</category>
      <category>多模态生成</category>
      <category>图像生成</category>
      <description>Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multim...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Haoyu Zhao, Zihao Zhao, Tianyu Deng et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 多模态生成, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> audio-visual generation, video generation, gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unified architecture raises a fundamental question: Can multimodal alignment improve the model&apos;s world reasoning, and what new evaluation paradigms do omni-modal inputs enable? To investigate this question, this work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning. Unlike existing evaluation frameworks for video generation and world models, which are often constrained by limited input modalities and evaluation settings where prompts closely match the target video content, our evaluation is specifically designed to exploit the multimodal inputs of Omni-Model. We construct a diverse set of novel tasks that require models to integrate complementary information across modalities. Specifically, we consider four scenarios, including implicit prompts paired with multiple frames, audio-image, prefix-videos, and audio-video inputs. Every single modality provides only partial evidence about the underlying event, requiring the model to jointly reason over the complementary semantic cues to infer latent event states and future dynamics. Across 517 evaluation instances, MiniMax-H3 achieves an overall success rate of 41.97%. Video-based Decision Reasoning yields the highest success rate at 56.00%, while Audio-based Disambiguation Reasoning is the weakest, reaching only 27.40%. These results indicate that effective multimodal integration remains key to fully exploiting the benefits of diverse input modalities. The project is available at https://github.com/gulucaptain/MiniMax-H3-Reason.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18323v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18323v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Visual Autoregressive Priors for RAW-to-sRGB Image Signal Processing</title>
      <link>https://arxiv.org/abs/2609.18302v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18302v1</guid>
      <pubDate>Wed, 16 Sep 2026 08:24:26 GMT</pubDate>
      <dc:creator>Tailai Chen, Xiaotong Luo, Yuan Gao et al.</dc:creator>
      <category>模型架构</category>
      <description>RAW-to-sRGB image signal processing (ISP) must recover perceptually faithful colors and fine details from sensor measurements, often under imperfect spatial alignment and missing camera metadata. This...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Visual Autoregressive Priors for RAW-to-sRGB Image Signal Processing</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tailai Chen, Xiaotong Luo, Yuan Gao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">RAW-to-sRGB image signal processing (ISP) must recover perceptually faithful colors and fine details from sensor measurements, often under imperfect spatial alignment and missing camera metadata. This paper presents, to the best of our knowledge, the first application of visual autoregressive (VAR) next-scale prediction over a discrete image codebook to the RAW-to-sRGB ISP task. We adapt a frozen 1.10\,B-parameter VAR backbone for RAW-conditioned ISP with only 32.93\,M trainable parameters (2.99\%), and propose a frequency-decomposed color loss that separately supervises low-frequency tone via wavelet LL cosine similarity and chromatic edges via detail-band $\ell_1$. On the Zurich RAW-to-sRGB benchmark, the method improves PSNR-Y from 21.31 to 21.89\,dB and reduces LPIPS from 0.276 to 0.218 on the full 1,204-image test set. Diagnostic experiments show that the VAR prior preserves structure well, but continuous color transfer remains the dominant bottleneck: oracle affine correction recovers 3.8\,dB, while learned color heads yield marginal gains.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18302v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18302v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM</title>
      <link>https://arxiv.org/abs/2609.18279v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18279v1</guid>
      <pubDate>Wed, 16 Sep 2026 08:02:37 GMT</pubDate>
      <dc:creator>Victor Bercy, Martyna Poreba, Michal Szczepanski et al.</dc:creator>
      <category>模型架构</category>
      <description>Vision Transformers (ViTs) have achieved state-of-the-art performance across a range of computer vision tasks, mainly thanks to the self-attention mechanism. However, its complexity, increasing quadra...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Victor Bercy, Martyna Poreba, Michal Szczepanski et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Vision Transformers (ViTs) have achieved state-of-the-art performance across a range of computer vision tasks, mainly thanks to the self-attention mechanism. However, its complexity, increasing quadratically with the number of tokens, remains the major obstacle to ViT efficiency and deployment at scale. Token merging reduces this cost by aggregating redundant tokens. Yet existing methods are typically evaluated within a single architecture, leaving open whether their effectiveness stems from the merging mechanism itself or from the specific decoder they are paired with. We extend Graph-Guided Token Merging (G2TM), a single module inserted early in a ViT-based network, beyond its original Segmenter setting. We evaluate G2TM across three semantic segmentation frameworks (Segmenter, SETR, EoMT) and three decoder families (Linear, Transformer-, convolution-based), as well as standard ViT image classification. Our results show that G2TM&apos;s behavior and accuracy-efficiency trade-off are consistent across every tested architecture for a given backbone size, indicating that its effectiveness is a property of the encoder rather than the decoder. G2TM also generalizes well to image classification, achieving an even smaller degradation in accuracy compared to semantic segmentation. We further find that G2TM&apos;s optimal hyperparameters, resulting in a consistent drop in GFLOPs of 22-47% and an increase in throughput by up to 74% for segmentation models on ADE20K dataset, depend primarily on the backbone&apos;s pre-training recipe and on the target dataset, rather than on the decoder choice.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18279v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18279v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] MS-RFD: Multi-Signal Release Frame Detection in Hammer Throw from Reconstructed 3D Trajectories</title>
      <link>https://arxiv.org/abs/2609.18260v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18260v1</guid>
      <pubDate>Wed, 16 Sep 2026 07:38:59 GMT</pubDate>
      <dc:creator>Ahmed Endris Hasen, Nikolaos Passalis, Tomi Vanttinen et al.</dc:creator>
      <category>模型架构</category>
      <description>Recent advances in artificial intelligence and computer vision are reshaping sports performance analysis by enabling automated detection, tracking, and performance analysis. In hammer throw, performan...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MS-RFD: Multi-Signal Release Frame Detection in Hammer Throw from Reconstructed 3D Trajectories</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ahmed Endris Hasen, Nikolaos Passalis, Tomi Vanttinen et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Recent advances in artificial intelligence and computer vision are reshaping sports performance analysis by enabling automated detection, tracking, and performance analysis. In hammer throw, performance is strongly determined by the kinematic conditions at release, particularly release speed, release angle, and release height. However, identifying the release instant from video typically requires manual frame-by-frame inspection, which is subjective and cumbersome in real-world training scenarios. In this paper, we present a fully automatic multi-signal release frame detection (MS-RFD) method for hammer throw using reconstructed 3D hammer trajectories. The proposed method integrates four complementary kinematic signals: speed dynamics, angular velocity transition, radial distance relative to the rotation center, and post-release trajectory linearity. These signals are fused to score and verify candidate release frames. MS-RFD is evaluated through the throwing-distance estimation error obtained from the release parameters estimated at the detected frame. An ablation study analyzes the contribution of each signal and compares alternative candidate selection strategies. The results show that speed dynamics and radial expansion provide the strongest signals for release frame detection, while angular velocity and post-release linearity provide smaller refinements.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18260v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18260v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models</title>
      <link>https://arxiv.org/abs/2609.18259v2</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18259v2</guid>
      <pubDate>Wed, 16 Sep 2026 07:38:41 GMT</pubDate>
      <dc:creator>Chunpu Xu, Zhixuan Liang, Yuhao Zhang et al.</dc:creator>
      <category>模型架构</category>
      <description>Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Chunpu Xu, Zhixuan Liang, Yuhao Zhang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This &quot;discretization bottleneck&quot; significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the ${M}^2$Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at https://github.com/cpaaax/M2Tok.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18259v2" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18259v2" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Evolving Error States: Failure-Aware Progressive Repair for Ultrasound Lesion Segmentation</title>
      <link>https://arxiv.org/abs/2609.18256v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18256v1</guid>
      <pubDate>Wed, 16 Sep 2026 07:35:19 GMT</pubDate>
      <dc:creator>Ziliang Wang, XuJiang Tang, Lu Yuting et al.</dc:creator>
      <category>模型架构</category>
      <description>Reliability under sparse and heterogeneous failures remains a fundamental challenge for medical image segmentation. High average accuracy can conceal a small set of structurally distinct and clinicall...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Evolving Error States: Failure-Aware Progressive Repair for Ultrasound Lesion Segmentation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ziliang Wang, XuJiang Tang, Lu Yuting et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Reliability under sparse and heterogeneous failures remains a fundamental challenge for medical image segmentation. High average accuracy can conceal a small set of structurally distinct and clinically consequential errors. Existing post-hoc correction methods alleviate this problem, but typically estimate false-positive and false-negative corrections from the same fixed prediction. This ignores the dynamic evolution of error states and limits the correction of complex cases. Inspired by iterative error feedback in structured prediction, we propose Failure-Aware Progressive Repair (FAPR). FAPR represents the current segmentation mask as a dynamic failure state and models each repair operation as a state-transition operator. Each accepted correction forms a new prediction state for subsequent error diagnosis and repair, enabling later operations to adapt to preceding changes. Conditional routing selectively activates necessary state transitions, while failure replay exposes the model to rare error states. By keeping the base segmentor frozen, FAPR preserves its established segmentation capability while improving difficult cases. Across three public ultrasound lesion segmentation benchmarks, FAPR improves mean DSC by 1.52%. On the very-hard subsets of BUSI and TN3K, the average gain reaches 13.77%.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18256v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18256v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Unified Response Geometry for Structured Pruning</title>
      <link>https://arxiv.org/abs/2609.18239v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18239v1</guid>
      <pubDate>Wed, 16 Sep 2026 07:15:09 GMT</pubDate>
      <dc:creator>Kaixiang Shu</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Structured pruning is commonly formulated as ranking individual channels, although channel responses can be complementary or cancel through downstream mixing. Motivated by these response interactions,...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Unified Response Geometry for Structured Pruning</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kaixiang Shu</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> imagen, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Structured pruning is commonly formulated as ranking individual channels, although channel responses can be complementary or cancel through downstream mixing. Motivated by these response interactions, we formulate pruning as the selection of a subset with large joint response capacity, followed by a separate functional realization step. Our unified response geometry maps each candidate set to \(M(D,R)=D^{1/2}RD^{1/2}\) and uses its determinant together with Schur-greedy residuals to select non-redundant coordinates. The same construction yields two information-conditioned instances: an unlabeled instance based on activation covariance, and a task-conditioned instance that combines activation and gradient variance for response scale with gradient correlation for complementarity. To convert the selected subset into an executable network, we fold predictable removed responses into successor weights through ridge compensation and recalibrate batch-normalization statistics, without fine-tuning the network. On ImageNet ResNet-50, the unlabeled instance reaches \(65.4\%\) and \(53.9\%\) Top-1 accuracy at 30\% and 40\% deletion, versus \(59.8\%\) and \(43.1\%\) for strength-only selection; the task-conditioned instance reaches \(67.7\%\) and \(56.3\%\) under the same protocol. A six-family screen shows architecture-dependent behavior, with positive relative contrasts in several convolutional and expansion-layer settings and clear boundary cases in windowed attention. These results support response geometry as a conditional principle for structured pruning, with its benefit determined jointly by the observed response and the architecture in which that response is realized.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18239v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18239v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] WISE: A Lightweight, Weakly-Supervised Model for Onboard Fire Smoke Detection and Localization</title>
      <link>https://arxiv.org/abs/2609.18227v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18227v1</guid>
      <pubDate>Wed, 16 Sep 2026 07:00:48 GMT</pubDate>
      <dc:creator>Sha Lu, Yu Sun, Liang Zhao et al.</dc:creator>
      <category>扩散模型</category>
      <description>Wildfire smoke detection from satellite imagery is critical for early warning and rapid response. For onboard satellite deployment, detection systems must operate under strict memory and latency const...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">WISE: A Lightweight, Weakly-Supervised Model for Onboard Fire Smoke Detection and Localization</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sha Lu, Yu Sun, Liang Zhao et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Wildfire smoke detection from satellite imagery is critical for early warning and rapid response. For onboard satellite deployment, detection systems must operate under strict memory and latency constraints while providing spatially informative outputs for downstream decision-making. Existing tile-level classification methods are computationally efficient but lack spatial localization, whereas pixel-level segmentation approaches provide detailed masks yet are typically too computationally demanding for real-time onboard execution. To address this gap, we propose WISE (Weakly-supervised Inference-efficient Smoke Extraction), a deployment-oriented framework for onboard fire smoke detection and localization. WISE leverages only tile-level annotations through a teacher-student distillation strategy, where an offline teacher provides soft spatial supervision to a lightweight WISE-Student optimized for efficient onboard inference. The student jointly predicts tile-level smoke presence and smoke probability maps within a single forward pass, enabling spatially informative detection under strict computational constraints. WISE was evaluated through in-orbit execution aboard the ISS-mounted IMAGIN-e payload. Three model variants achieve average inference times of 0.10 s, 0.14 s, and 0.26 s per tile, indicating near-real-time per-tile inference within onboard resource limits. Ground-based experiments on Landsat 5 and Landsat 8 imagery further indicate effective detection and spatially informative localization. The best-performing variant achieves a mean tile-level F1 score of 0.964 and a mean pixel-level F1 score of 0.750 across 10 runs, while containing only 0.12M parameters and requiring approximately 3 GFLOPs. Together, these results indicate that WISE is a practical candidate for low-latency wildfire smoke monitoring from space under onboard resource constraints.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18227v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18227v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] A Lightweight CNN Integrated Compact Convolutional Transformer for Multi-Scale Feature Learning and reducing computational complexity for breast cancer mammography image detection and classification</title>
      <link>https://arxiv.org/abs/2609.18212v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18212v1</guid>
      <pubDate>Wed, 16 Sep 2026 06:42:38 GMT</pubDate>
      <dc:creator>Md Taimur Ahad, Ainuddin Ahmed</dc:creator>
      <category>模型架构</category>
      <description>Over the years, Convolutional Neural Networks (CNNs) have demonstrated strong capability in cancer detection and classification using medical images. However, CNN-based models often struggle to captur...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Lightweight CNN Integrated Compact Convolutional Transformer for Multi-Scale Feature Learning and reducing computational complexity for breast cancer mammography image detection and classification</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Md Taimur Ahad, Ainuddin Ahmed</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Over the years, Convolutional Neural Networks (CNNs) have demonstrated strong capability in cancer detection and classification using medical images. However, CNN-based models often struggle to capture long-range contextual dependencies. In such scenarios, integrating Compact Convolutional Transformer (CCT) architectures after the CCT layer allows CNN-extracted features to reshape into compact patch tokens using a CCT tokenizer, followed by the addition of positional embeddings to preserve spatial structure. Using 5-fold cross-validation, the model was tested on 3 sets of breast cancer mammography. With only 250,435 parameters, the model achieved 99%-100% accuracy across 3 datasets, indicating robust generalization. Explainable AI (XAI) was integrated into the model to explain the breast cancer classification process to enhance clinical trust. The results indicate that the proposed framework is suitable for computer-aided diagnosis systems, particularly in resource-constrained clinical environments. The novelty of the proposed CNN-integrated CCT overcomes the limitation of CNN&apos;s gradient degradation in the last layers by integrating convolutional tokenization with transformer-based learning. Lighter than ViT, which is effective in capturing long-range dependencies, the model has also proven efficient in breast cancer classification by capturing long-range dependencies among breast tissue regions.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18212v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18212v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs</title>
      <link>https://arxiv.org/abs/2609.18210v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18210v1</guid>
      <pubDate>Wed, 16 Sep 2026 06:41:36 GMT</pubDate>
      <dc:creator>Yuhang Zhu, Meiyi Zhu, Yunkai Dang et al.</dc:creator>
      <category>多模态生成</category>
      <description>UAV-borne imaging has advanced from megapixel to gigapixel sensors, shifting aerial perception from recognizing individual targets to understanding entire dynamic scenes. We characterize this demand a...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yuhang Zhu, Meiyi Zhu, Yunkai Dang et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">UAV-borne imaging has advanced from megapixel to gigapixel sensors, shifting aerial perception from recognizing individual targets to understanding entire dynamic scenes. We characterize this demand as Wide-area Spatio-temporal Scene Understanding (WSTU), which requires wide-area coverage, per-target resolution, and temporal continuity at once, a combination existing datasets lack. To fill this gap, we introduce an ultra-High-resolution (12768x9564) Airborne Remote-sensing Dataset (HARD) annotated at three levels for object detection, multi-object tracking, and scene-level visual question answering. Ultra-high-resolution imagery raises per-frame processing time to seconds. At that scale latency can no longer be ignored in evaluation. Thus, we propose a latency-aware metric for multi-object tracking called streaming-HOTA (s-HOTA). Extensive baseline experiments show how ultra-high-resolution processing reshapes each task. For detection, the end-to-end pipeline affects accuracy and speed as much as the detector itself does. For tracking, high latency charges the association axis far more unevenly than the detection axis, and association is where pipelines diverge. As a result, the pipeline that performs best offline can lose its lead under s-HOTA. For VQA, vision-language models remain weak at cross-frame identity binding and cannot transfer their single-frame gains to it. Together these findings show that the baselines we evaluate fall short of WSTU. HARD provides the data and the systematic baselines to advance it.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18210v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18210v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Stealthy in Semantics, Antagonistic in Space: Attacking Visible-Infrared Object Detectors via Object-Level Misalignment</title>
      <link>https://arxiv.org/abs/2609.18133v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18133v1</guid>
      <pubDate>Wed, 16 Sep 2026 05:16:19 GMT</pubDate>
      <dc:creator>Yueqi Zhu, Qi Ming, Guo Cheng et al.</dc:creator>
      <category>模型架构</category>
      <description>Visible-infrared object detectors are used for robust perception under challenging illumination and weather conditions. Current physical attacks apply conspicuous patches to spatially aligned target r...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Stealthy in Semantics, Antagonistic in Space: Attacking Visible-Infrared Object Detectors via Object-Level Misalignment</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Yueqi Zhu, Qi Ming, Guo Cheng et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Visible-infrared object detectors are used for robust perception under challenging illumination and weather conditions. Current physical attacks apply conspicuous patches to spatially aligned target regions, which are noticeable to human observers. Meanwhile, most of these methods only perturb the appearance within the aligned region, without explicitly targeting the correspondence between modalities or the fusion process. In this paper, we propose CamoShift, an adversarial framework for visible-infrared object detection. By combining visual camouflage with object-level infrared shifting, CamoShift breaks cross-modal spatial alignment and disrupts fusion. Specifically, the Semantic Camouflage Module (SCM) generates a stealthy camouflaged patch that can be attached to the host object and maintains its effectiveness in the infrared branch through an RGB-IR adapter. The Object-level Spatial Decoupling Module (OSDM) shifts the infrared target evidence in a scale-aware manner, so as to break object-level correspondence and disrupt cross-modal fusion. Then, the Harmonic Adversarial loss (HarAdv loss) further balances attack strength and visual stealth during optimization. To the best of our knowledge, we are the first to target both visual stealthiness and attack success in visible-infrared object detection. Extensive experimental results show that CamoShift achieves a superior balance between attack effectiveness and visual stealth. Code and models will be available on GitHub.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18133v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18133v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] MCLC-NET: Multimodal Continual Learning for Leaf Counting</title>
      <link>https://arxiv.org/abs/2609.18129v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18129v1</guid>
      <pubDate>Wed, 16 Sep 2026 05:07:52 GMT</pubDate>
      <dc:creator>Ruchi Bhatt, Pratibha Kumari, Shreya Bansal et al.</dc:creator>
      <category>模型架构</category>
      <description>Leaf counting is an important task in plant phenotyping for monitoring plant growth and estimating crop yield. Most existing methods rely on RGB images, but their performance is often affected by occl...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">MCLC-NET: Multimodal Continual Learning for Leaf Counting</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Ruchi Bhatt, Pratibha Kumari, Shreya Bansal et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Leaf counting is an important task in plant phenotyping for monitoring plant growth and estimating crop yield. Most existing methods rely on RGB images, but their performance is often affected by occlusion, lighting variations, and other real-world challenges. Additional modalities, such as depth and thermal images, can provide useful complementary information. However, multimodal leaf counting remains underexplored. Also, many existing methods assume that all training data are available simultaneously, which is impractical in real agricultural settings, where data is collected over time from multiple sources. To address these challenges, we propose MCLC-NET, a multimodal continual learning framework for leaf counting. It learns tasks sequentially using a memory-based strategy with a memory buffer to retain important samples from previous tasks. We also introduce MMLC, a real-world multimodal leaf-counting dataset designed for a domain incremental scenario (DIS) in CL. It contains RGB, depth, and thermal images collected across different crop types under varying environmental conditions, arranged in three orderings: crop-wise, time-wise, and mixed. Experimental results, averaged over three random seeds, demonstrate that MCLC-NET consistently outperforms existing methods across all three task orderings, achieving the lowest AMSE of 0.675$\pm$0.027, 0.542$\pm$0.069, and 0.745$\pm$0.057, respectively.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18129v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18129v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] PRISM: Predictive Representation of Interaction Style and Motion for Social Robot Navigation</title>
      <link>https://arxiv.org/abs/2609.18125v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18125v1</guid>
      <pubDate>Wed, 16 Sep 2026 05:01:27 GMT</pubDate>
      <dc:creator>Bo-Han Chen, Hiromu Taketsugu, Norimichi Ukita</dc:creator>
      <category>模型架构</category>
      <description>Humans often observe others before interacting and adjust their behavior accordingly. Robot navigation in crowds, however, often represents pedestrians mainly by observed geometric states, leaving ind...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">PRISM: Predictive Representation of Interaction Style and Motion for Social Robot Navigation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Bo-Han Chen, Hiromu Taketsugu, Norimichi Ukita</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Humans often observe others before interacting and adjust their behavior accordingly. Robot navigation in crowds, however, often represents pedestrians mainly by observed geometric states, leaving individual differences in interaction tendencies implicit. We propose PRISM (Predictive Representation of Interaction Style and Motion), a framework that infers interaction traits from passive observations of human-human interactions. PRISM encodes human trajectories into a continuous ordinal latent space with a transformer encoder trained by Rank-N-Contrast loss, and pairs each inferred trait with a temporal-stability score supplied to the navigation policy. In randomized crowd simulations, PRISM reduces collision rates over the geometry-only baseline and yields small improvements in navigation-time and path-length metrics. These results suggest the utility of passive latent-trait inference for social navigation in dynamic crowds.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18125v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18125v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] [图像生成] A Comprehensive Review of Generative Physical Artificial Intelligence</title>
      <link>https://arxiv.org/abs/2609.18111v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18111v1</guid>
      <pubDate>Wed, 16 Sep 2026 04:29:50 GMT</pubDate>
      <dc:creator>Satyam Gaba, Krutiksinh Rana, Siva Sai et al.</dc:creator>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelligence (GPAI). These agentic AI...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Comprehensive Review of Generative Physical Artificial Intelligence</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Satyam Gaba, Krutiksinh Rana, Siva Sai et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, diffusion model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelligence (GPAI). These agentic AI systems autonomously perceive, reason, and act in complex real-world situations. This survey comprehensively analyzes GPAI systems, focusing on their architectural foundations, current applications, and key limitations. We introduce a taxonomy of five distinct approaches: Robot Foundation Models (RFMs) for cross-platform skill transfer; Vision-Language Action (VLA) models for end-to-end multi-modal perception and control; Large Behavior Models (LBMs) for human-like movement generation; Diffusion Policy Models (DPMs) for diffusion model-based temporally coherent action generation; and World Foundation Models (WFMs) for physics-compliant simulation and data generation. We examine how these approaches complement each other: WFMs generate training data for VLAs and DPMs, RFMs enable cross-platform deployment of learned policies, while LBMs provide motion priors for natural behavior. Through examples across autonomous vehicles, industrial automation, healthcare robotics, and humanoid systems, we identify significant performance improvements and summarize promising research directions in data-efficient learning, sim-to-real transfer, edge-compatible architectures, and safety frameworks. These insights advance embodied AI for IoT-connected environments where intelligent agents interact with networked sensors, actuators, and edge devices.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18111v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18111v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception</title>
      <link>https://arxiv.org/abs/2609.18100v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18100v1</guid>
      <pubDate>Wed, 16 Sep 2026 04:08:51 GMT</pubDate>
      <dc:creator>Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Synthetic data can reduce the cost of collecting and annotating training data for robotic perception, but generating sensor observations that preserve the characteristics relevant to downstream percep...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> generative adversarial network, gan, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Synthetic data can reduce the cost of collecting and annotating training data for robotic perception, but generating sensor observations that preserve the characteristics relevant to downstream perception remains challenging, particularly for sonar imagery. In this work, we investigate whether conventional image-fidelity metrics adequately reflect the downstream perception performance of GAN-generated synthetic sonar data. We employ a Pix2Pix conditional generative adversarial network with four discriminator configurations characterized by different receptive fields: PixelGAN, PatchGAN-16, PatchGAN-70, and ImageGAN. The models are trained using sonar imagery from two datasets and evaluated using conventional image-fidelity metrics, including Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Mean Squared Error (MSE). To complement these pixel-level measures with task-oriented evaluation, YOLOX-S, YOLOX-L, and Faster R-CNN detectors are trained exclusively on real sonar imagery and subsequently evaluated on the GAN-generated images using identical test samples and annotations across all discriminator configurations. The results reveal a discrepancy between image-fidelity and downstream object-detection performance: the configuration achieving the best SSIM, PSNR, and MSE does not consistently yield the best detection performance. In particular, PatchGAN configurations achieve strong downstream detection results despite not achieving the highest pixel-level similarity scores. These findings suggest, for the datasets and models considered, pixel-level image-fidelity metrics alone may not consistently capture the task-relevant realism of synthetic sonar observations and motivate the use of task-aware evaluation for synthetic sensor data intended for robotic perception.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18100v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18100v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Mask 2D-3D: Adaptive Dual-Masked Autoencoder Network for Image-to-Point Cloud Registration</title>
      <link>https://arxiv.org/abs/2609.18088v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18088v1</guid>
      <pubDate>Wed, 16 Sep 2026 03:47:07 GMT</pubDate>
      <dc:creator>Zhixin Cheng, Jiacheng Deng, Xiaotian Yin et al.</dc:creator>
      <category>模型架构</category>
      <description>Detection-free methods for image-to-point cloud registration are prone to erroneous correspondences caused by domain and modality discrepancies, limited sensitivity of feature extractors, and the pres...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Mask 2D-3D: Adaptive Dual-Masked Autoencoder Network for Image-to-Point Cloud Registration</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhixin Cheng, Jiacheng Deng, Xiaotian Yin et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, mae, masked autoencoder</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Detection-free methods for image-to-point cloud registration are prone to erroneous correspondences caused by domain and modality discrepancies, limited sensitivity of feature extractors, and the presence of non-overlapping regions. The Masked Autoencoder (MAE) has shown strong performance in visual representation for images and point clouds. It may be helpful to apply this approach to image-to-point cloud registration, a task that requires unified feature extraction and accurate cross-modal correspondences. Standard MAE&apos;s random masking may overlook key regions due to limited camera views, reducing registration effectiveness. To address this, we propose the Intermodal Dual-MAE Framework (ID-MAE) with a Similarity-based RL Masking Strategy (SRLM), which adaptively masks informative positions by leveraging cross-modal similarity and reinforcement learning, thus narrowing the modality gap. Our method enhances cross-modal representation learning by enforcing representation consistency during feature extraction, thereby enabling more reliable 2D-3D correspondence estimation. Experiments on RGB-D Scenes v2 and 7-Scenes benchmarks show that our method achieves state-of-the-art performance in image-to-point cloud registration.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18088v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18088v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models</title>
      <link>https://arxiv.org/abs/2609.18084v2</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18084v2</guid>
      <pubDate>Wed, 16 Sep 2026 03:44:51 GMT</pubDate>
      <dc:creator>Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev et al.</dc:creator>
      <category>图像生成</category>
      <description>Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every region requires equ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every region requires equal adjustment. This paper tests that assumption on five architecturally diverse VLAs (OpenVLA-OFT, $π_0$, SmolVLA, DTP, Octo; 93M-7B parameters). Measuring per-region adaptation cost as normalized parameter displacement under region-isolated fine-tuning reveals an adaptation spectrum in which appearance shifts concentrate cost in the vision encoder, instruction shifts in the language backbone, and novel-object shifts in the vision encoder together with the action head, across all five architectures. To exploit this structure, we introduce a pipeline that observes, diagnoses, allocates, and adapts. From ten unlabeled target observations and without fine-tuning, the diagnostic estimates per-region cost by combining reference-free gradient and Monte Carlo Dropout signals with a Centered Kernel Alignment score against a cached source reference; the allocator converts the estimates into variable-rank LoRA adapters under a parameter budget and freezes well-calibrated regions; and standard LoRA fine-tuning trains the resulting adapters. The diagnostic ranks regions within each deployment at a median Spearman of 0.91, and the allocation matches or exceeds uniform LoRA at every budget we tested on LIBERO and CALVIN. On a physical xArm-7, the pipeline matches full fine-tuning under an instruction-wording shift with 0.04% of its trainable parameters, and on five held-out scenes evaluated without retraining it leads every baseline, with 11-23 successes of 30 rollouts against 8-18 for the strongest parameter-efficient baseline at equal or larger budgets and 2-11 for full fine-tuning. These results suggest that adaptation cost in VLAs is structured enough to measure before fine-tuning begins.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18084v2" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18084v2" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[视频生成] [模型架构] [扩散模型] vidax: A Unified JAX Framework for Video Generative Models on Accelerator Meshes</title>
      <link>https://arxiv.org/abs/2609.18077v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18077v1</guid>
      <pubDate>Wed, 16 Sep 2026 03:26:45 GMT</pubDate>
      <dc:creator>Congyue Deng</dc:creator>
      <category>视频生成</category>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Open-source video generative models ship almost exclusively as PyTorch/CUDA reference implementations. This leaves Cloud TPU pods without a production-ready inference path, despite offering large, cos...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">vidax: A Unified JAX Framework for Video Generative Models on Accelerator Meshes</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Congyue Deng</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 视频生成, 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> diffusion, video generation, vae, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Open-source video generative models ship almost exclusively as PyTorch/CUDA reference implementations. This leaves Cloud TPU pods without a production-ready inference path, despite offering large, cost-effective accelerator memory pools ideal for long-sequence spatiotemporal attention. We present vidax, an open-source JAX/Flax inference engine and zero-copy PyTorch-to-JAX weight translator for modern video generation architectures. vidax covers a diverse set of spatiotemporal models --- including Diffusion Transformers, omnimodal Mixture-of-Transformers, 3D VAEs, text encoders, and native samplers --- with zero PyTorch dependency in the execution path. The framework unifies 1D tensor parallelism with DeepSpeed-Ulysses sequence parallelism on a single JAX sharding mesh, integrates TPU flash-attention kernels, and implements per-layer weight offloading to support reference resolutions that exceed single-device memory. We benchmark compile times, latency, and peak memory utilization on TPU v4-8 hardware, and document real-world numerical bugs surfaced during checkpoint translation. vidax is released open-source as a baseline for JAX and TPU video generation research.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18077v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18077v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 视频生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Finder: Agentic Closed-Loop Object Finding for Embodied Grounding</title>
      <link>https://arxiv.org/abs/2609.18058v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18058v1</guid>
      <pubDate>Wed, 16 Sep 2026 03:02:45 GMT</pubDate>
      <dc:creator>Shixiong Xu, Zhiyuan Chen, Song Ding et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Finding the object referred to by language in a partially observed 3D scene is a core capability for embodied agents. Existing approaches either couple object search with online exploration, which can...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Finder: Agentic Closed-Loop Object Finding for Embodied Grounding</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shixiong Xu, Zhiyuan Chen, Song Ding et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Finding the object referred to by language in a partially observed 3D scene is a core capability for embodied agents. Existing approaches either couple object search with online exploration, which can be costly when relevant observations have already been captured, or query pre-built open-vocabulary maps and scene graphs in a static, one-shot fashion. We present Finder, an agentic closed-loop object-finding primitive for embodied grounding. Instead of treating grounding as passive retrieval from a fixed scene representation, Finder maintains a typed loop state that links query-conditioned planning, scoped evidence gathering, candidate verification, and accept/continue/abort control. When evidence is incomplete or ambiguous, the loop can redirect subsequent perception and comparison rather than simply returning the top retrieved object. On open-vocabulary embodied Object Retrieval in Habitat/HM3D and real-world RGB-D scenes, Finder improves the averaged 1m success rate by 15.75 points over strong baselines. The same primitive also transfers to sequential object grounding and embodied object-centric question answering, improving spatial and temporal localization without changing the inner grounding protocol. Project page: https://finder-vln.github.io.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18058v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18058v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] Position Anchor Tuning: Towards Efficient Adaptation of Pre-Trained Point Cloud Transformers</title>
      <link>https://arxiv.org/abs/2609.18056v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18056v1</guid>
      <pubDate>Wed, 16 Sep 2026 02:59:29 GMT</pubDate>
      <dc:creator>Zheng Liu, Xin Gao, Jinchao Zhu et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Parameter-efficient fine-tuning (PEFT) has recently emerged as a pivotal research direction for adapting pre-trained point cloud transformers to diverse downstream tasks. Although existing methods ach...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Position Anchor Tuning: Towards Efficient Adaptation of Pre-Trained Point Cloud Transformers</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zheng Liu, Xin Gao, Jinchao Zhu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Parameter-efficient fine-tuning (PEFT) has recently emerged as a pivotal research direction for adapting pre-trained point cloud transformers to diverse downstream tasks. Although existing methods achieve excellent fine-tuning performance with high parameter efficiency, they ignore inference efficiency. To tackle this problem, a novel PEFT method termed position anchor tuning (PAT) is proposed in this paper. As multi-head attention (MHA) and feed-forward network (FFN) are computation-heavy blocks in pre-trained transformers, PAT decreases their computational cost through token aggregation-expansion pairs. Each pair comprises a token aggregation module (TAM) and a token expansion module (TEM). For MHA and FFN blocks, TAMs extract representative tokens from their input tokens based on position anchors in 3D space. These extracted tokens, rather than the original input tokens, are processed by the blocks, thereby reducing the number of tokens involved in computation. Then, TEMs propagate the learned representations back to the original input tokens. Since TAMs are solely responsible for capturing task-specific representations, base-sharing low-rank adaptation (BSLoRA) is further introduced to enable them to learn such representations effectively with only a small number of trainable parameters. Extensive experiments on widely used benchmarks demonstrate that PAT performs comparably to state-of-the-art methods while incurring significantly lower computational overhead and fewer trainable parameters.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18056v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18056v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [图像生成] SetPlanner: A Lightweight Plug-in Point-Set Planner for Frozen SAM</title>
      <link>https://arxiv.org/abs/2609.18037v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18037v1</guid>
      <pubDate>Wed, 16 Sep 2026 02:38:59 GMT</pubDate>
      <dc:creator>Dawei Yan, Yuezhe Yang, Menglan Ruan et al.</dc:creator>
      <category>模型架构</category>
      <category>图像生成</category>
      <description>Segment Anything Models provide reusable priors, yet they require user prompts and cannot support fully automatic instrument segmentation. Automatic prompting is difficult for thin, articulated, refle...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">SetPlanner: A Lightweight Plug-in Point-Set Planner for Frozen SAM</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Dawei Yan, Yuezhe Yang, Menglan Ruan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> lora, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Segment Anything Models provide reusable priors, yet they require user prompts and cannot support fully automatic instrument segmentation. Automatic prompting is difficult for thin, articulated, reflective, and partly occluded tools, where several configurations can be valid. We formulate automatic prompting as lightweight point-set planning and isolate the point source under a frozen pathway. To this end, we present SetPlanner, a 1.52M-parameter plug-in point-set planner for frozen SAM. The plug-in preserves SAM&apos;s point-prompt interface and enables reuse across backbones. SetPlanner plans complete unordered K-point sets from geometry-aware targets with a permutation-aware conditional flow. SAM decodes eight candidates; their consensus readout yields a ground-truth-free prediction. Across three endoscopic datasets, SetPlanner wins all six transfer routes over a LoRA-adapted system. Under our frozen-pathway protocol, SetPlanner reaches 0.934 Dice on Kvasir-Instrument and recovers 96% of a 44.4-point localization gap, while candidate disagreement ranks low-Dice cases at AUROC 0.969.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18037v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18037v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[扩散模型] [图像生成] Newer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations</title>
      <link>https://arxiv.org/abs/2609.18007v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.18007v1</guid>
      <pubDate>Wed, 16 Sep 2026 01:54:12 GMT</pubDate>
      <dc:creator>Shesh Narayan Gupta, Nik Bear Brown</dc:creator>
      <category>扩散模型</category>
      <category>图像生成</category>
      <description>Text-to-image generative models are widely used in professional and creative settings, yet how they represent gender across occupations -- and whether newer models are fairer -- remains poorly underst...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Newer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Shesh Narayan Gupta, Nik Bear Brown</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 扩散模型, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> stable diffusion, lora, diffusion, diffusion model, text-to-image</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Text-to-image generative models are widely used in professional and creative settings, yet how they represent gender across occupations -- and whether newer models are fairer -- remains poorly understood across multiple generations. We evaluate gender representation across 20 occupations, 5 prompt templates, and 4 Stable Diffusion model generations (SD 1.5, SD 2.1, SDXL, SD 3 Medium), generating 8,000 images with n = 100 per occupation-model cell (5 prompts x 20 images), and classifying all with DeepFace. Across the 8,000 open-source images, 76.4% show male subjects (95% CI [75.1%, 78.7%], p &lt; 2.2 x 10^-16, Benjamini-Hochberg adjusted). More strikingly, 57.6% of images for historically female-coded occupations show male subjects (raw p = 3.43 x 10^-22, BH-adjusted p = 1.71 x 10^-21). All nine significant tests reported in this paper survive BH correction across 10 tests. When compared against U.S. Bureau of Labor Statistics workforce data, models underrepresent women by 20-46pp on average, with particularly large deviations for near gender-balanced occupations: scientist (48% female in BLS, 82-99% male in model outputs) and cleaner (46% female in BLS, 80-92% male in outputs). Model generations do not improve steadily: bias worsens from SD 1.5 to SDXL before partially recovering in SD 3 Medium. A preliminary comparison with GPT-image-1 on five occupations suggests lower bias than open-source models, though the practical effect is small (Cramer&apos;s V = 0.080) and the comparison is exploratory. No model achieves gender parity.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.18007v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.18007v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 扩散模型, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing</title>
      <link>https://arxiv.org/abs/2609.17953v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.17953v1</guid>
      <pubDate>Wed, 16 Sep 2026 00:25:31 GMT</pubDate>
      <dc:creator>Sihao Ding, Santosh Vasa, Aditi Ramadwar et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Vision-Language Models (VLMs) can produce Natural Language Explanations (NLEs) that sound plausible yet remain inconsistent with the visual evidence they cite. We present Explanation-Driven Counterfac...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Sihao Ding, Santosh Vasa, Aditi Ramadwar et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-16</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Vision-Language Models (VLMs) can produce Natural Language Explanations (NLEs) that sound plausible yet remain inconsistent with the visual evidence they cite. We present Explanation-Driven Counterfactual Testing (EDCT), an intervention-based protocol that extracts visual concepts cited in a model&apos;s explanation, applies verified minimal edits to them, and tests whether the resulting answer and explanation remain consistent with the edited image. Using this protocol, we create EDCT-Bench, a comprehensive benchmark spanning three complementary domains: knowledge-intensive visual question answering (OK-VQA), safety-critical driving (DriveLM), and 3D spatial reasoning (3DSRBench). Across the evaluated VLMs, EDCT reveals substantial faithfulness gaps, with models frequently producing responses inconsistent with verified visual changes. Finally, our fine-tuning study suggests that EDCT-generated counterfactuals provide high-impact training signals.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.17953v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.17953v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Rapid Loss of the Sierra Nevada&apos;s Largest Trees Driven by Fire</title>
      <link>https://arxiv.org/abs/2609.17925v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.17925v1</guid>
      <pubDate>Tue, 15 Sep 2026 23:33:59 GMT</pubDate>
      <dc:creator>Fabien H. Wagner, Dan J. Dixon, Christopher W. Woodall et al.</dc:creator>
      <category>模型架构</category>
      <description>Large trees disproportionately contribute to biomass storage, habitat structure, and ecosystem functioning. However, their distribution and health dynamics remain poorly quantified at a regional scale...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Rapid Loss of the Sierra Nevada&apos;s Largest Trees Driven by Fire</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Fabien H. Wagner, Dan J. Dixon, Christopher W. Woodall et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> u-net</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Large trees disproportionately contribute to biomass storage, habitat structure, and ecosystem functioning. However, their distribution and health dynamics remain poorly quantified at a regional scale. Here, a deep learning model (U-Net-ID) and canopy height models derived from sub-meter aerial imagery from 2020 were used to delineate all individual trees with crown area $\geq$ 100 m$^2$ across the Sierra Nevada Floristic Province. The model was trained using more than 3.3 million synthetic tree crowns and achieved a median Intersection over Union (IoU) of 0.602 when validated against an independent dataset of 20,273 crowns. A total of 6,515,705 large trees were mapped, occurring across approximately 78.7% of the Sierra Nevada Floristic Province. The spatial distribution of large trees showed associations with elevation, temperature, and precipitation. Using Sentinel-2 time series from 2020 to 2025, tree health dynamics were characterized by extracting spectral trajectories for each crown and applying BFAST breakpoint detection algorithm combined with a disturbance classification framework to identify mortality, disturbance, and recovery trajectories of individual trees. Wildfires, estimated from CAL FIRE fire perimeters, were identified as the dominant driver of large-tree mortality, killing 10% of all large trees in the Sierra Nevada, with mortality strongly concentrated during the extreme 2020-2021 fire seasons.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.17925v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.17925v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control</title>
      <link>https://arxiv.org/abs/2609.17909v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.17909v1</guid>
      <pubDate>Tue, 15 Sep 2026 23:05:39 GMT</pubDate>
      <dc:creator>Mingyang Chen, Shengdong Chen, Xiaoxiao Fu et al.</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint key...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Mingyang Chen, Shengdong Chen, Xiaoxiao Fu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trained on connected multi-prompt videos to supervise a block-level causal student through distribution-matching distillation; and (3) Low-cost real-time interaction, combining four-step generation with context-preserving streaming to support 832 x 480 inference at 24 FPS at an estimated server rental cost of approximately USD 0.009 per stream-minute. Zing-0.5 achieves an overall score of 81.0 and a consistency score of 88.5 across 158 WBench Navigation cases. A joint-control demonstration shows a text-directed event change during continued navigation without restarting generation. We release the model weights, inference code, and Zing-SGLang serving implementation to support further work on playable generated worlds.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.17909v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.17909v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[图像生成] ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software</title>
      <link>https://arxiv.org/abs/2609.17885v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.17885v1</guid>
      <pubDate>Tue, 15 Sep 2026 22:23:17 GMT</pubDate>
      <dc:creator>Kratika Bhagtani, Kusha Sridhar, Maziyar Baran Pouyan et al.</dc:creator>
      <category>图像生成</category>
      <description>Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks. Enterprise Resource Planning (...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kratika Bhagtani, Kusha Sridhar, Maziyar Baran Pouyan et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks. Enterprise Resource Planning (ERP) systems run the finance, procurement, inventory, and customer operations of organizations worldwide, and pose distinct challenges for computer-use agents: dense interfaces, coordinated multi-step interactions, and errors that alter persistent business records rather than surfacing on screen. Existing enterprise benchmarks rely on proprietary platforms or on simulated approximations of such software. We introduce ERPBench, a benchmark that evaluates screenshot-only agents on a live and reproducible ERP system and scores each task against ground-truth values in its database. Beyond the benchmark, we present a production-grade harness that gates agent actions behind human approval for safe deployment, which ERPBench runs autonomously. Evaluating six closed and open-source agents, we demonstrate that strong general GUI performance does not transfer to enterprise reliability. Even when an agent reaches the right form and saves it, the stored record is often wrong: some agents save in up to 85% of runs but write the correct value in as few as 3%. We further characterize failure modes specific to enterprise workflows.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.17885v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.17885v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] Can VLMs Reliably Assess Sidewalk Accessibility Attributes from Pedestrian-Level Imagery?</title>
      <link>https://arxiv.org/abs/2609.17882v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.17882v1</guid>
      <pubDate>Tue, 15 Sep 2026 22:15:04 GMT</pubDate>
      <dc:creator>Seung Jae Lieu, Diego Morra, Chiara Cadoni et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>An important component of urban accessibility, particularly for wheelchair users and people with reduced mobility, is sidewalk compliance with measurable requirements. We test whether effective width,...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Can VLMs Reliably Assess Sidewalk Accessibility Attributes from Pedestrian-Level Imagery?</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Seung Jae Lieu, Diego Morra, Chiara Cadoni et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">An important component of urban accessibility, particularly for wheelchair users and people with reduced mobility, is sidewalk compliance with measurable requirements. We test whether effective width, longitudinal slope, cross slope, and pavement condition can be assessed reliably from pedestrian-level imagery using vision-language models (VLMs). We present the first application of sampling-based conformal prediction (CP) for VLM-based accessibility assessment. We evaluate four VLMs on 514 sidewalk images from Seoul, South Korea, with field-measured ground truth. Conformal calibration attains the nominal 90% coverage for all models and attributes, but the calibrated regions differ in informativeness. Effective width yields the most informative estimates, with a mean interval half-width of about 1.0 m for the best model. Since every model overestimates width, asymmetric calibration shortens the intervals by up to 33% at unchanged coverage. Longitudinal slope is marginally informative, cross-slope intervals are too wide to resolve regulatory thresholds, and pavement-condition sets degenerate to all five grades (A-E) for three of the four models. Uncalibrated intervals from raw sampling dispersion cover only 17-47% of field-measured values at a nominal 90% level. Among the images with the most self-consistent responses, these intervals miss the field-measured value in up to 96% of cases. Response self-consistency is therefore not evidence of accuracy, and sampling dispersion cannot be interpreted as uncertainty until it has been calibrated against field-measured ground truth. No quantitative attribute reaches the precision required for general compliance assessment, but CP identifies from calibration data alone which attributes can support screening of segments far from the thresholds. We release the annotated pedestrian-level images and their corresponding field-measured attribute values.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.17882v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.17882v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [图像生成] Lumen: Parameter-Efficient Alignment of Pretrained Vision and Language Encoders for Zero-Shot Computational Pathology</title>
      <link>https://arxiv.org/abs/2609.17868v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.17868v1</guid>
      <pubDate>Tue, 15 Sep 2026 21:47:58 GMT</pubDate>
      <dc:creator>Kiarash Tajbakhsh, Abdelrahman Faqieh, Michael Jopiti et al.</dc:creator>
      <category>多模态生成</category>
      <category>图像生成</category>
      <description>Pathology vision-language models are commonly built by pretraining or fine-tuning large encoders on paired image-caption data. We asked whether a pathology vision-language model can instead be assembl...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Lumen: Parameter-Efficient Alignment of Pretrained Vision and Language Encoders for Zero-Shot Computational Pathology</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Kiarash Tajbakhsh, Abdelrahman Faqieh, Michael Jopiti et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 图像生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, gan</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Pathology vision-language models are commonly built by pretraining or fine-tuning large encoders on paired image-caption data. We asked whether a pathology vision-language model can instead be assembled by parameter-efficient alignment of frozen unimodal foundation models, leaving their pretrained representations untouched. Here we present Lumen, which aligns frozen Virchow2 and BioMedBERT backbones using rank-4 adapters and projection heads, training only 0.40% of the total parameters on the public QUILT-1M corpus. Across nine public zero-shot patch benchmarks, Lumen achieved the highest mean chance-corrected balanced accuracy, 0.546 versus 0.461 for the strongest baseline (paired difference 0.086, 95% CI 0.042-0.136). On lymph-node metastasis detection, Lumen reached an AUROC of 0.964 (95% CI 0.956-0.971) on 4,214 held-out internal slides and 0.955 (95% CI 0.942-0.966) on 2,368 slides across nine external cohorts and six organs. At the internally calibrated threshold, it outperformed all vision-language baselines, with a balanced accuracy of 0.909 (95% CI 0.896-0.923) internally and 0.915 (95% CI 0.902-0.929) externally. Lumen performed competitively across the evaluations, with the exception of cross-modal retrieval, where it ranked third behind CONCH and PathGen-L/14. Fully fine-tuning both encoders gave Lumen no consistent benefit over low-rank adaptation, although it improved retrieval. Aligning frozen unimodal foundation models therefore yields strong and transferable performance at patch and slide level while training only a small fraction of the parameters.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.17868v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.17868v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 图像生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] Mechanism-Level Evaluation for Vision-Language Models: Controlled Activation-Replacement Diagnosis of Gender Bias</title>
      <link>https://arxiv.org/abs/2609.16651v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.16651v1</guid>
      <pubDate>Tue, 15 Sep 2026 05:10:41 GMT</pubDate>
      <dc:creator>Zhipeng Zhao, Wenxu Wang, Peishun Liu et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Behavioral benchmarking reveals \emph{what} biases exist in vision-language models but not \emph{which internal components} are most sensitive to targeted intervention, precluding principled intervent...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Mechanism-Level Evaluation for Vision-Language Models: Controlled Activation-Replacement Diagnosis of Gender Bias</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhipeng Zhao, Wenxu Wang, Peishun Liu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, vision-language model</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Behavioral benchmarking reveals \emph{what} biases exist in vision-language models but not \emph{which internal components} are most sensitive to targeted intervention, precluding principled intervention. We argue for mechanism-level evaluation as a necessary complement, demonstrating causal mediation analysis as a diagnostic instrument for gender bias. We decompose gender-cue effects into controlled indirect effects attributable to specific-layer activations and direct effects through all other pathways, producing layer-by-layer mechanistic signatures. Across six models spanning three architectural families (LLaVA-1.5, LLaVA-NeXT, InstructBLIP at 7B/13B) and two 8B-scale architectures, three findings emerge: language-layer activations exhibit the greatest output sensitivity under controlled intervention, with the direct component often carrying the opposite sign; architectural choices redistribute layer-wise sensitivity to activation replacement; and counterfactual scores diverge from surface-level scores, exposing implicit associations. Systematic ablation validates internal consistency. An intervention experiment finds that the average indirect effect (AIE) and downstream intervention effectiveness are only weakly correlated (Pearson $r = 0.33$), and the layer with the second-largest AIE produces near-zero bias change---indicating that mechanistic diagnosis captures activation-replacement sensitivity but does not, by itself, identify optimal intervention targets. These results show mechanism-level evaluation captures architecture-specific sensitivity patterns that behavioral benchmarks cannot; pairing both should become standard NLP practice. Code: https://github.com/zhaozhipeng1997/CARD-GenderBias.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.16651v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.16651v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] [模型架构] ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models</title>
      <link>https://arxiv.org/abs/2609.16647v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.16647v1</guid>
      <pubDate>Tue, 15 Sep 2026 05:08:34 GMT</pubDate>
      <dc:creator>Zhipeng Zhao, Zhaoqiang Wei, Peishun Liu et al.</dc:creator>
      <category>多模态生成</category>
      <category>模型架构</category>
      <description>Gender bias in large vision-language models (LVLMs) undermines their fairness and reliability, compromising output trustworthiness. Current mitigation methods rely on training-phase adjustments or pos...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhipeng Zhao, Zhaoqiang Wei, Peishun Liu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成, 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Gender bias in large vision-language models (LVLMs) undermines their fairness and reliability, compromising output trustworthiness. Current mitigation methods rely on training-phase adjustments or post-hoc calibration, but face limitations in dynamic visual bias mitigation. These include inability to capture real-time visual-textual incongruence, dependence on predefined gender bias taxonomies, and degraded cross-modal alignment with emergent bias patterns. To address these challenges, we propose ViD, a causally-inspired framework that analyzes attention mechanisms across five distinct patterns, revealing confounding effects from strong language priors. ViD demonstrates that visual-to-language cross-attention effectively suppresses bias while preserving general reasoning capabilities and text generation quality. ViD incorporates dual mechanisms: backdoor adjustment counters strong language priors, while refined token selection in decoding layers optimizes processing. This enhances model robustness and inference efficiency. Our integrated approach significantly mitigates gender bias across multidimensional social attributes in LVLMs, improving visual grounding and output fairness. Cross-benchmark validation shows ViD reduces gender bias by 14.7\% on single-attribute evaluations (FACET) and achieves significant improvements on image captioning tasks (MS COCO), with gender bias score improving from 0.6708 to 0.9978 for LLaVA. Crucially, these improvements require no additional training overhead, making ViD a scalable and practical solution for bias mitigation in LVLMs.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.16647v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.16647v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成, 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] What Do Hallucinations Reveal About Multimodal Reasoning? Diagnosing Visual Grounding Failures via Contrastive Decoding Probes</title>
      <link>https://arxiv.org/abs/2609.16646v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.16646v1</guid>
      <pubDate>Tue, 15 Sep 2026 05:07:01 GMT</pubDate>
      <dc:creator>Zhipeng Zhao, Wenxu Wang, Peishun Liu et al.</dc:creator>
      <category>多模态生成</category>
      <description>When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instruments for understanding behavior. We address this by ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">What Do Hallucinations Reveal About Multimodal Reasoning? Diagnosing Visual Grounding Failures via Contrastive Decoding Probes</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Zhipeng Zhao, Wenxu Wang, Peishun Liu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instruments for understanding behavior. We address this by asking: can we use large vision-language models (LVLMs) as experimental instruments for studying their own failure dynamics? Focusing on visual hallucination, we introduce SAFE, a training-free decoding framework that contrasts visually-grounded and vision-ablated generation paths to produce a token-level contrastive grounding score that identifies when the model favors linguistic priors over visual evidence. This signal serves dual roles: as a practical proxy for detecting visually-ungrounded tokens, and as the basis for decoding-time penalties. Our analysis yields three empirical observations: visual dependency decays over generation, hallucinations co-occur in temporal clusters, and early intervention reduces clustering without substantially degrading fluency. On MMHalBench, SAFE substantially outperforms all compared baselines; results elsewhere are more mixed. We argue that designing contrastive probes exemplifies a broader mission: using models as instruments for scientific understanding. Code: https://github.com/zhaozhipeng1997/SAFE_public.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.16646v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.16646v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Multimodal Emergency Vehicle Classification via Audio-Visual Transformers and Knowledge Distillation</title>
      <link>https://arxiv.org/abs/2609.16535v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.16535v1</guid>
      <pubDate>Tue, 15 Sep 2026 02:34:07 GMT</pubDate>
      <dc:creator>Vijay John, Amar Dabaja</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Emergency vehicle detection in autonomous driving is a safety-critical perception task that demands robustness under diverse and adverse real-world conditions. Existing approaches rely on a single mod...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Multimodal Emergency Vehicle Classification via Audio-Visual Transformers and Knowledge Distillation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Vijay John, Amar Dabaja</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer, distillation</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-15</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Emergency vehicle detection in autonomous driving is a safety-critical perception task that demands robustness under diverse and adverse real-world conditions. Existing approaches rely on a single modality, either audio or video, which leads to systematic failure when that modality is degraded: microphone-based systems fail in noisy urban environments, and camera-based systems fail at night or under occlusion. This report presents AVNet, a multimodal audio-visual transformer that classifies emergency vehicles (ambulance, fire engine, police car) and road background using both audio and video, while gracefully handling the absence of either modality at inference time. AVNet introduces three key contributions: (1) a temporally aligned cross-modal fusion module that performs second-level cross-attention between audio spectrogram tokens and video frame tokens, exploiting their exact temporal correspondence without any learned alignment mechanism; (2) learned null embeddings that substitute for missing modality tokens, enabling a single unified model to operate in audio-only, video-only, or joint audio-visual mode without retraining; and (3) a knowledge distillation training strategy in which specialist unimodal teacher models transfer inter-class dark knowledge into the multimodal student fusion branch via soft probability targets. Evaluated on 281 clips from the Google AudioSet dataset, AVNet achieves 66.6% overall accuracy in audio-visual mode, outperforming the audio-only branch by +10.4% and the video-only branch by +15.0%. The largest per-class gain is observed for the hardest class, Ambulance, where fusion achieves +29.5% over either unimodal branch alone, demonstrating that the two modalities provide complementary information that the aligned cross attention mechanism successfully exploits.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.16535v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.16535v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Tracing the Origins: Legacy Codec Identification in Neural Audio Transcoding</title>
      <link>https://arxiv.org/abs/2609.14916v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.14916v1</guid>
      <pubDate>Mon, 14 Sep 2026 02:08:15 GMT</pubDate>
      <dc:creator>Wonje Heo, Shinee Youn, Yooshin Kim et al.</dc:creator>
      <category>模型架构</category>
      <description>Residual Vector Quantization (RVQ)-based neural audio codecs (NACs) enable high-fidelity audio distribution at unprecedentedly low bitrates through discrete token-based representations. However, this ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Tracing the Origins: Legacy Codec Identification in Neural Audio Transcoding</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Wonje Heo, Shinee Youn, Yooshin Kim et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> dit, transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-14</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Residual Vector Quantization (RVQ)-based neural audio codecs (NACs) enable high-fidelity audio distribution at unprecedentedly low bitrates through discrete token-based representations. However, this shift disrupts traditional forensics, as non-linear neural transcoding obscures the underlying traces of legacy compression. This study defines the forensic gap and proposes a Transformer-based framework designed to leverage the hierarchical and temporal dependencies inherent in RVQ sequences. By modeling inter-layer causal relationships and dynamic forensic significance, our model effectively disentangles superimposed artifacts from legacy-to-neural transcoding. Experimental results achieve 97%+ accuracy for codec identification and robust joint identification performance across 32-128 kbps. These results demonstrate that traditional codec traces persist even after neural transcoding, supporting the feasibility and necessity of neural-codec-aware audio forensics.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.14916v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.14916v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Speak to the City: Multimodal Resolution for Outside-the-Vehicle References</title>
      <link>https://arxiv.org/abs/2609.14691v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.14691v1</guid>
      <pubDate>Sun, 13 Sep 2026 17:48:22 GMT</pubDate>
      <dc:creator>Alireza Parchami, Artin Saberpour, Robin Connor Schramm et al.</dc:creator>
      <category>模型架构</category>
      <description>As autonomous vehicles and Extended Reality (XR) headsets enable novel in-car interactions, seamlessly querying physical landmarks, known as Outside-the-Vehicle Referencing (OVR), remains challenging ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Speak to the City: Multimodal Resolution for Outside-the-Vehicle References</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Alireza Parchami, Artin Saberpour, Robin Connor Schramm et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> transformer</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-13</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">As autonomous vehicles and Extended Reality (XR) headsets enable novel in-car interactions, seamlessly querying physical landmarks, known as Outside-the-Vehicle Referencing (OVR), remains challenging due to ego-motion and referential ambiguity. We present a robust, multimodal OVR framework fusing user gaze and natural language to identify Points of Interest (POIs). To address the scarcity of dynamic vehicular data, we developed a VR-based pipeline synchronizing 360-degree transit videos with vehicle GNSS telemetry. Through a user study (N=46) mapping passenger head orientation into a 3D geospatial Digital Twin, we captured authentic gaze-speech behaviors. We subsequently trained a lightweight Transformer network, leveraging LLMs to dynamically align continuous spatial gaze vectors with discrete verbal context. Experimental results demonstrate high accuracy and low computational overhead, achieving an 83.33% Top-1 accuracy (87.72% Top-2) and an average inference time of 24.3 milliseconds. This real-time paradigm effectively resolves referential ambiguity, enabling context-aware spatial retrieval for passengers within the vehicle.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.14691v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.14691v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] Grounded in Sound: Reinforcement Learning with a Frozen Acoustic Judge to Curb ASR Insertion Hallucinations</title>
      <link>https://arxiv.org/abs/2609.14455v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.14455v1</guid>
      <pubDate>Sun, 13 Sep 2026 11:50:52 GMT</pubDate>
      <dc:creator>Tingzhen Xiong, Rilin Chen, Weiwei Li et al.</dc:creator>
      <category>模型架构</category>
      <description>When reinforcement learning (RL) is used for post-training automatic speech recognition (ASR), the reward almost always lives in the text space: it compares a hypothesis with the reference and never c...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Grounded in Sound: Reinforcement Learning with a Frozen Acoustic Judge to Curb ASR Insertion Hallucinations</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Tingzhen Xiong, Rilin Chen, Weiwei Li et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vit, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-13</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">When reinforcement learning (RL) is used for post-training automatic speech recognition (ASR), the reward almost always lives in the text space: it compares a hypothesis with the reference and never checks whether the hypothesis is supported by the audio. On highly regular speech this licenses a shortcut - guessing from a strong language prior rather than listening. Once the acoustics degrade, the shortcut runs unchecked and emits fluent but ungrounded words, i.e., insertion errors. We propose an acoustic-fidelity reward: a GRPO reward augmented with a separately pretrained, permanently frozen, non-autoregressive character-level wav2vec2-CTC acoustic judge, used strictly at training and absent at inference, where a single model decodes greedily. Trained on LibriSpeech and evaluated across a six-tier difficulty gradient including real AMI meeting speech (33,282 utterance-condition instances), the method reduces insertion errors by 28.3% on close-talking AMI-IHM and 22.3% on far-field AMI-SDM, while lowering WER on AMI-SDM from 35.89% to 34.71% and showing no detectable WER difference on the other five tiers, against a schedule-matched WER-GRPO baseline. The insertion reduction holds under a meeting-level clustered bootstrap. Four prespecified analyses support content-conditioned insertion calibration: output collapses 85-90% on unintelligible audio that preserves energy and voice activity; the gain is not recovered by the evaluated 32-best CTC rescoring configuration, yet RL internalizes it into a single greedy decoding run; and policy-only confidence yields lower insertion-AURC in all four evaluated settings. We frame this as a mechanism paper, demonstrated in one instantiation: a 7B speech LLM with a 0.3B CTC judge.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.14455v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.14455v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[模型架构] [扩散模型] Conditional Quantum Flow Matching for Data-Scarce Physiological Signal Augmentation</title>
      <link>https://arxiv.org/abs/2609.14019v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.14019v1</guid>
      <pubDate>Sat, 12 Sep 2026 16:06:56 GMT</pubDate>
      <dc:creator>Chi-Sheng Chen, Samuel Yen-Chi Chen</dc:creator>
      <category>模型架构</category>
      <category>扩散模型</category>
      <description>Generative augmentation is a standard remedy for label scarcity in physiological signal classification, but existing quantum generative models start from uninformative noise, ignoring class structure ...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">Conditional Quantum Flow Matching for Data-Scarce Physiological Signal Augmentation</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Chi-Sheng Chen, Samuel Yen-Chi Chen</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 模型架构, 扩散模型</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> ddpm, flow matching, dit</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-12</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Generative augmentation is a standard remedy for label scarcity in physiological signal classification, but existing quantum generative models start from uninformative noise, ignoring class structure that is already available. We propose Conditional Quantum Flow Matching (CQFM): a single 306-parameter circuit, conditioned on both flow time and class label, transports a compact class-conditional prior toward the target distribution. Quantum flow matching as published is unconditional, so this is to our knowledge the first conditional one, and the first EEG augmentation on a parameterized quantum circuit. A nonnegative spectral embedding removes the need for tomography at readout. On BCI Competition IV-2a, starting from a prior rather than noise is worth $+5.1$ accuracy points over QuDDPM (9/9 subjects), though at that operating point a class-conditional Gaussian matches CQFM. Where the prior fails the transport earns its keep: given one transferred from other subjects it regains $+7.2$ TSTR points (9/9).</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.14019v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.14019v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 模型架构, 扩散模型 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
    <item>
      <title>[多模态生成] A Low-Latency Interactive System for Real-Time Video Understanding Based on VLMs</title>
      <link>https://arxiv.org/abs/2609.13986v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.13986v1</guid>
      <pubDate>Sat, 12 Sep 2026 14:59:35 GMT</pubDate>
      <dc:creator>Punan Dai, Jun Xu, Bingcong Lu et al.</dc:creator>
      <category>多模态生成</category>
      <description>Vision-language models are extending video understanding from offline clip analysis to continuous interactive streaming, but most research still emphasizes model capability rather than deployable low-...</description>
      <content:encoded><![CDATA[
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; max-width: 800px; margin: 0 auto; padding: 20px;">
  <h2 style="color: #1a1a1a; margin-bottom: 10px;">A Low-Latency Interactive System for Real-Time Video Understanding Based on VLMs</h2>
  
  <div style="background: #f5f5f5; padding: 15px; border-radius: 8px; margin: 20px 0;">
    <p style="margin: 5px 0;"><strong>👤 作者：</strong> Punan Dai, Jun Xu, Bingcong Lu et al.</p>
    <p style="margin: 5px 0;"><strong>📂 分类：</strong> 多模态生成</p>
    <p style="margin: 5px 0;"><strong>🔖 关键词：</strong> vision-language model, vlm</p>
    <p style="margin: 5px 0;"><strong>📅 发布：</strong> 2026-09-12</p>
  </div>
  
  <div style="background: white; padding: 20px; border-left: 4px solid #4CAF50; margin: 20px 0;">
    <h3 style="color: #4CAF50; margin-top: 0;">📝 摘要</h3>
    <p style="line-height: 1.8; color: #333;">Vision-language models are extending video understanding from offline clip analysis to continuous interactive streaming, but most research still emphasizes model capability rather than deployable low-latency interaction. This paper presents a unified edge-cloud system for real-time video VLM applications. Lightweight phone, smart glasses, PC, and pseudo-replay clients publish video and speech to a server runtime that provides shared ASR/TTS, session orchestration, backend adaptation, response delivery, and archive-backed measurement. The system integrates six representative video VLM backends with streaming or interaction-oriented capabilities and evaluates them across backend runtime, media transport, client-observed latency, and interaction behavior. With suitable backend selection and the WebRTC path, the tested system reaches approximately 0.9 to 1.0 s to first VLM text and 1.3 to 1.5 s to first non-silent TTS audio, while exposing backend adaptation costs and differences in real-time interaction behavior.</p>
  </div>
  
  <div style="margin-top: 30px; padding: 15px; background: #e3f2fd; border-radius: 8px;">
    <p style="margin: 5px 0;"><strong>🔗 链接：</strong></p>
    <p style="margin: 10px 0;">
      <a href="https://arxiv.org/abs/2609.13986v1" style="color: #2196F3; text-decoration: none; margin-right: 20px;">📄 arXiv 页面</a>
      <a href="https://arxiv.org/pdf/2609.13986v1" style="color: #FF5722; text-decoration: none;">📥 PDF 下载</a>
    </p>
  </div>
  
  <div style="margin-top: 20px; padding: 10px; background: #fff3e0; border-radius: 8px; font-size: 0.9em;">
    <p style="margin: 5px 0; color: #666;">💡 <strong>提示：</strong>这是一篇 AIGC 领域的最新论文，涵盖 多模态生成 等主题。</p>
  </div>
</div>
]]></content:encoded>
      <enclosure url="https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png" type="image/png"/>
    </item>
  </channel>
</rss>