← 返回名录条目信息来自其公开主页与公开发布内容。按本站规范,页面不展示任何联系方式。 需要更正或删除?通过收录与更正通道提交,24 小时内处理。
L
Lian
LLM inference engineer in training · vLLM internals · KV cache · prefix caching · CUDA/Triton (learning)
- 公司
- —
- 位置
- China
- Stars
- 0
- 粉丝 / 仓库
- 4
代表作品 / 项目
- qwen3-prefix-cache-bench⭐ 0Reproducible Qwen3-1.7B prefix cache benchmark on RTX 3070 Laptop (8GB). Hand-written reference inference loop + vLLM v1 comparison.
- derLogik⭐ 0Profile README
- tiny-qwen⭐ 0A minimal PyTorch re-implementation of Qwen 3.5
- vllm⭐ 0A high-throughput and memory-efficient inference and serving engine for LLMs
- multi-model-agent⭐ 0Multi-agent collaboration system using LangGraph + FastAPI + Next.js (engineering demo, not inference work).