Zhean Xu
|
891d57b4db
|
Add various optimizations and Mega MoE benchmarks (#316)
* Merge with private repo
* Add Mega MoE Benchmark
* Minor fix
* Update
---------
Co-authored-by: Chenggang Zhao <chenggangz@deepseek.com>
|
2026-04-24 18:41:37 +08:00 |
|
Chenggang Zhao
|
7f2a703ed5
|
[Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes (#304)
* Merge with private repo
* Update README
* Update README
* Update README
* Add PyTorch requirements
* Fix sync scopes for MQA logits (#256)
* Update README
|
2026-04-17 09:45:14 +08:00 |
|
Zhean Xu
|
0f5f266202
|
Multiple updates and refactorings (#280)
|
2026-01-16 17:06:52 +08:00 |
|
Ray Wang
|
38f8ef73a4
|
Multiple updates and refactorings (#231)
|
2025-11-21 17:49:47 +08:00 |
|
Simon Mo
|
59f2c07cf2
|
Add SM100 kernels (#201)
Signed-off-by: simon-mo <simon.mo@hey.com>
|
2025-09-29 17:07:28 +08:00 |
|
Chenggang Zhao
|
80ceeb2c76
|
Add SM90 kernels (#200)
|
2025-09-29 17:00:23 +08:00 |
|
Ray Wang
|
3f71de7aa9
|
Make various updates and fixes (#198)
|
2025-09-25 16:19:07 +08:00 |
|