Enable scaled FP8 (e4m3fn) KV cache on ROCm (AMD GPU) (#3290)

Co-authored-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: AdrianAbeyta <Adrian.Abeyta@amd.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: root <root@gt-pla-u18-08.pla.dcgpu>
Co-authored-by: mawong-amd <156021403+mawong-amd@users.noreply.github.com>
Co-authored-by: ttbachyinsda <ttbachyinsda@outlook.com>
Co-authored-by: guofangze <guofangze@kuaishou.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
Co-authored-by: jacobthebanana <50071502+jacobthebanana@users.noreply.github.com>
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>

This commit is contained in:

Adrian Abeyta

2024-04-03 16:15:55 -05:00

committed by

GitHub

parent 3dcb3e8b98

commit 2ff767b513

41 changed files with 2592 additions and 142 deletions

1

.gitignore vendored

View File

@@ -181,6 +181,7 @@ _build/
 # hip files generated by PyTorch
 *.hip
 *_hip*
 hip_compat.h
 # Benchmark dataset
 *.json

Enable scaled FP8 (e4m3fn) KV cache on ROCm (AMD GPU) (#3290)

1 .gitignore vendored Unescape Escape View File

1

.gitignore vendored

View File