Logo
Explore Help
Register Sign In
Serendipity/CTI-Inference-Opt
1
1
Fork 0
You've already forked CTI-Inference-Opt
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
Files
69d49cd282d5607ef0872a9229d5743fa51f2a17
CTI-Inference-Opt/代码/code
T
History
OwnerSunshine530 69d49cd282 revert: MoE加权+attention输出布局两刀(评测净负35.85>34.64,大中间张量/跨步写代价>省的clone)。保留消同步刀单独测
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 20:56:27 +08:00
..
tests
perf: _triton_block_meta 消除最后一个host同步(grid用shape派生上界,空block在kernel内mask空跑)
2026-06-19 20:51:37 +08:00
bench.py
feat: 真稀疏MoE(capacity分组,只算top-k,cutlass baddbmm,无host同步)
2026-06-17 21:05:55 +08:00
build_env.sh
fix: build_env.sh 简化为纯净版本(避免 CUDA 预热导致异常)
2026-06-12 21:55:09 +08:00
EXPERIMENTS.md
docs: 收尾 — 最终67.998/记录RepEncoder预计算尝试与结论
2026-06-16 13:18:48 +08:00
infer.py
revert: MoE加权+attention输出布局两刀(评测净负35.85>34.64,大中间张量/跨步写代价>省的clone)。保留消同步刀单独测
2026-06-19 20:56:27 +08:00
requirements.txt
revert: requirements.txt 还原为原始完整依赖列表
2026-06-12 21:24:22 +08:00
RISKS.md
docs: 潜在风险说明(RepEncoder预计算合规灰区/max_feasign一致性)与合规保底
2026-06-15 20:44:57 +08:00
Powered by Gitea Version: 26.3.1 Page: 1010ms Template: 30ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API