← 返回名录条目信息来自其公开主页与公开发布内容。按本站规范,页面不展示任何联系方式。 需要更正或删除?通过收录与更正通道提交,24 小时内处理。
H
Huadong Wu
PhD student in Beijing, chasing models that see, hear, and actually reason. Multimodal / speech / vision LLMs. Mostly PyTorch.
- 公司
- Institute of Automation, CAS
- 位置
- Beijing, China
- Stars
- 3
- 粉丝 / 仓库
- 3
代表作品 / 项目
- tinyvlm⭐ 1A minimal, hackable vision-language model. ~500 lines of readable PyTorch — train on one GPU, modify without fighting abstractions.
- visual-token-soup⭐ 1Adaptive visual-token pruning for efficient VLM inference — monkey-patch drop-in, 2-4x faster LLaVA.
- mm-reason-bench⭐ 1Fine-grained probes for multimodal reasoning: where vision-language models actually break.
- verl-omni⭐ 0Multimodal RL training framework for diffusion & omni models
- DataFlow⭐ 0Easy Data Preparation with latest LLMs-based Operators and Pipelines.