Chenggang Zhao
|
7f2a703ed5
|
[Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes (#304)
* Merge with private repo
* Update README
* Update README
* Update README
* Add PyTorch requirements
* Fix sync scopes for MQA logits (#256)
* Update README
|
2026-04-17 09:45:14 +08:00 |
|
Zhean Xu
|
0f5f266202
|
Multiple updates and refactorings (#280)
|
2026-01-16 17:06:52 +08:00 |
|
Ray Wang
|
3f71de7aa9
|
Make various updates and fixes (#198)
|
2025-09-25 16:19:07 +08:00 |
|
Ray Wang
|
d9c363f86f
|
Make various updates and fixes:
- Add support for legacy CUDA versions; now compatible with CUDA 12.3 and newer
- Add support for NVRTC compilation
- Other fixes and code refactoring
|
2025-08-02 19:52:22 -07:00 |
|
Ray Wang
|
9da4a23561
|
Add more GPU architectures support (#112)
* Add more GPU architectures support
* Update layout.py
* Optimize performance, Add SM90 support, Add 1D2D SM100 support
* Add fmtlib submodule at commit 553ec11
---------
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
|
2025-07-18 11:32:22 +08:00 |
|