ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

state-space-model

2 posts tagged with "state-space-model"

July 8, 2026

Video Understanding: Through A Temporal Lens(面向时间维度的视频理解)

Thong Thanh Nguyen 的博士论文,从时间维度系统性地推进视频理解,围绕三大研究目标(推进边界 / 高效捕获时间 / 有效捕获时间)提出五项贡献:(1) MAMA 用减性角度间隔对比 + 元优化样本加权自动标注海量视频;(2) READ 在轻量适配器瓶颈内嵌 RNN,仅微调适配器即在低资源时间定位/视频摘要上超越全量微调;(3) GSMT 用门控状态空间模型(SSM)对密集采样帧建全局语义,配套 Ego-QA/MAD-QA 长视频问答基准;(4) MCL 用运动感知对比学习 + 强运动 tube 选择做时间全景场景图;(5) MSCL 多尺度跨尺度对比 + 特征金字塔做时间定位;并给出面向时间的 LVLM 训练配方(QueryFormer 接口 + 时间训练方案 + 记忆库 + MoE 四步扩展)。

video-understanding vision-language long-context contrastive-learning parameter-efficient state-space-model
July 7, 2026

LinGen: 线性复杂度的高分辨率分钟级文本到视频生成

LinGen 用线性复杂度的 MATE 模块(Bidirectional Mamba2 + TESA)替换 DiT 中二次复杂度的自注意力,首次实现单 GPU 上高分辨率分钟级视频生成,最高 15× FLOPs 加速。

diffusion-model video-generation efficient-attention state-space-model

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.