Files
nvfp4-megamoe-kernel/src/nvfp4_megamoe_kernel/__pycache__/symm_buffer.cpython-312.pyc

37 lines
3.4 KiB
Plaintext
Raw Normal View History

<EFBFBD>
X<>j <00> <00><><00>dZddlZddlZddlmZeejjdd<04><00>Z Gd<05>d<06>Z
deded ed
ed ed e
f d <0A>Z y)z<>Symmetric buffer for NVLink cross-rank all-reduce in mega_moe.
Replaces deep_gemm.mega.SymmBuffer and get_symm_buffer_for_nvfp4_mega_moe.
API matches the DeepGEMM signature used in the vLLM deepseek_v4.py patch.
<EFBFBD>N<>MEGA_MOE_DEBUG<55>0c<00><00>eZdZdZd<02>Zy)<04>
SymmBuffera<EFBFBD>Symmetric NVLink buffer for expert-parallel cross-rank communication.
Matches the DeepGEMM SymmBuffer interface expected by the vLLM patch:
- .x: staged activation (FP4 packed)
- .x_sf: staged activation scales (UE4M3 packed)
- .topk_idx: top-k expert indices
- .topk_weights: top-k routing weights
- .buffer: underlying CUDA buffer
- .group: process group
c <00>p<00>||_||_||_||_||_||_t jj<00>}|dz}|dz} t j||dzt j|<07><03>|_ t j||t j|<07><03>|_ t j||t j|<07><03>|_t j||t j |<07><03>|_t j||t j$|<07><03>|_t(rt+d|jj,<00>d|jj,<00>d|jj,<00>d|j"j,<00>d|j&j,<00><00>
<EFBFBD>yy) N<>@<00>)<02>dtype<70>devicez[SymmBuffer] x=z x_sf=z
topk_idx=z topk_weights=z buffer=)<17>group<75> num_experts<74>max_num_tokens<6E>top_k<5F> hidden_size<7A>intermediate_size<7A>torch<63>cuda<64>current_device<63>empty<74>int8<74>x<>uint32<33>x_sf<73>int32<33>topk_idx<64>float32<33> topk_weights<74>bfloat16<31>bufferr<00>print<6E>shape)
<EFBFBD>selfr r rrrrr <00>sf_k_groups_hidden<65>sf_k_groups_inters
<20>E/home/openclaw/dev/nvfp4-mojo/src/nvfp4_megamoe_kernel/symm_buffer.py<70>__init__zSymmBuffer.__init__st<00><00><1A><04>
<EFBFBD>&<26><04><18>,<2C><04><1B><1A><04>
<EFBFBD>&<26><04><18>!2<><04><1E><16><1A><1A>*<2A>*<2A>,<2C><06>)<29>V<EFBFBD>4<><1A>-<2D>&<26>9<><19><17><1B><1B> <1A>K<EFBFBD>1<EFBFBD>,<2C><17>*<2A>*<2A>V<EFBFBD>
<EFBFBD><04><06><1A>K<EFBFBD>K<EFBFBD> <1A>.<2E><17>,<2C>,<2C>v<EFBFBD>
<EFBFBD><04> <09><1E> <0B> <0B> <1A>E<EFBFBD><17>+<2B>+<2B>f<EFBFBD>
<EFBFBD><04> <0A>"<22>K<EFBFBD>K<EFBFBD> <1A>E<EFBFBD><17>-<2D>-<2D><06>
<EFBFBD><04><19> <1C>k<EFBFBD>k<EFBFBD> <1A>K<EFBFBD><17>.<2E>.<2E><16>
<EFBFBD><04> <0B>
<1A> <11>O<EFBFBD>D<EFBFBD>F<EFBFBD>F<EFBFBD>L<EFBFBD>L<EFBFBD>><3E><16><04> <09> <09><0F><0F>7H<37>I<1E>"<22>m<EFBFBD>m<EFBFBD>1<>1<>2<>.<2E><14>AR<41>AR<41>AX<41>AX<41>@Y<>Z<1C> <20>K<EFBFBD>K<EFBFBD>-<2D>-<2D>.<2E>0<> 1<> <1A>N)<05>__name__<5F>
__module__<EFBFBD> __qualname__<5F>__doc__r&<00>r'r%rrs <00><00> <08>*1r'rr rrrr<00>returnc<00>"<00>t||||||<05>S)z<>Allocate a symmetric buffer for the NVFP4 mega_moe kernel.
API matches deep_gemm.mega.get_symm_buffer_for_nvfp4_mega_moe.
)r)r r rrrrs r%<00>"get_symm_buffer_for_nvfp4_mega_moer/Gs <00><00> <16> <0A>{<7B>N<EFBFBD>E<EFBFBD><13>&<26> <06>r') r+<00>osr<00>torch.distributed<65> distributed<65>dist<73>int<6E>environ<6F>getrrr/r,r'r%<00><module>r7sy<00><01><04> 
<EFBFBD> <0C> <20><14>R<EFBFBD>Z<EFBFBD>Z<EFBFBD>^<5E>^<5E>$4<>c<EFBFBD>:<3A>;<3B><0E>61<>61<>r<06><14><06><18><06> <0F> <06>
<15> <06> <1B> <06><10>r'