01
量化:从 Scale 到 W4A8 的完整坐标
ai-systems / llm-inference
llm-inferencequantizationprefilldecode
+1
02
从 AVX/AMX 到 Tensor Core:CPU 工程师理解 LLM 量化
ai-systems / llm-inference
llm-inferencequantizationcpugpu
03
FP4/FP8 量化:值域、Scale 与运行时合同
ai-systems / llm-inference
llm-inferencequantizationgpu
04
AWP Profiling API
toolbox
profilingapigpucpu
+2
05
Gavel: Heterogeneity-Aware Cluster Scheduling (OSDI'20)
ai-systems / distributed-training
schedulingclusterheterogeneousGPU
+1
06
GPU Trace 时间分解与通信计算重叠分析
ai-systems / profiling
GPUProfilingPerformanceDistributed Training
07
CUDA Agent
ai-systems / gpu-computing
GPUCUDARLLLM
+1
08
Compute-bound vs Memory-bound:推理的两大瓶颈
ai-systems / llm-inference
LLMInferencePerformanceGPU
+3
09
HTA 算法原理与实现
ai-systems / profiling
profilingpytorchgpudistributed-training
+2
10
Critical Path of AI Trace
ai-systems / profiling
AITraceCritical PathGPU
+1
11
PTX 技术详解
ai-systems / gpu-computing
cudagpuptxsass
+1
12
SAC: Sharing-Aware Caching in Multi-Chip GPUs
ai-systems / gpu-computing
GPUCacheMulti-ChipArchitecture
+1
13
GPU Communication
ai-systems / gpu-computing
GPUAI
14
GPU Architecture Deep Dive
ai-systems / gpu-computing
GPUCUDAParallel ComputingAI Infrastructure