← 返回名录条目信息来自其公开主页与公开发布内容。按本站规范,页面不展示任何联系方式。 需要更正或删除?通过收录与更正通道提交,24 小时内处理。
Y
Yuxuan Chen
CS grad @ ZJU. Multimodal & speech models. Trying to make machines see, hear, and talk.
- 公司
- Zhejiang University
- 位置
- Hangzhou, China
- Stars
- 3
- 粉丝 / 仓库
- 0
代表作品 / 项目
- LiteSpeechLM⭐ 1Lightweight speech language model built from scratch. Treats speech as discrete tokens and models them with a small transformer.
- OmniAlign⭐ 1Modular toolkit for aligning vision, audio, and speech encoders to LLMs via trainable projectors
- speech-tokenizer-arena⭐ 1Unified benchmark for evaluating and comparing neural speech tokenizers (EnCodec, SpeechTokenizer, DAC, HuBERT)
- dbmask⭐ 0Discover, mask, and validate sensitive data in SQL databases — deterministic masking with built-in per-row verification. Works across PostgreSQL, MySQL, SQL Server, Oracle, and SQLite via SQLAlchemy.
- verl-omni⭐ 0Multimodal RL training framework for diffusion & omni models