* ggml-et: Add performance logging * ggml-et: Quants helpers * ggml-et: Add MUL_MAT kernel * ggml-et: Add ROPE kernel * ggml-et: Add RMS_NORM kernel * ggml-et: Add GLU kernel * ggml-et: Add SOFT_MAX kernel * ggml-et: Add GET_ROWS kernel * ggml-et: Add CONT kernel * ggml-et: Add SET_ROWS kernel * ggml-et: Add MUL_MAT_ID kernel * ggml-et: Build et kernels as part of ggml * ggml-et: Embed kernels with fs fallback * ggml-et: Build fixes * ggml-et: Add MUL_MAT F32xF32 op * ggml_et: Add MUL_MAT_ID op * ggml-et: Disable offloading for debug * ggml-et: Refactor out block ops * ggml-et: ggml backend API changes * ggml-et: Add RESHAPE/TRANSPOSE to supported * ggml-et: Add CONT_F16 * ggml-et: Add supported ops doc * gglm-et: Initial doc * ggml-et: Remove runtime import hacks We can now import the runtime by a simple find_package(), so we can cleanup the CMakeLists.txt. * ggml-et: Fix GET_ROWS kernel Fix lost batch dimension. Also clean vibe-comments. * ggml-et: Fix SET_ROWS kernel Remove incorrect broadcasting guard. * ggml-et: Use custom instruction for fp32->fp16 * ggml-et: Vectorize set_rows fp32->fp16 * ggml-et: Fix ROPE kernel (yarn) ggml-et: fix et_logf WIP: Fix ramp WIP: fix ROPE! * ggml-et: Better sinf * ggml-et: Fix SOFT_MAX Add `max_bias` and `sink` support. * ggml-et: Fix CONT Reorder from contiguous write to read with atomic stores. * ggml-et: Fix elmap kernel Remainder handlin * ggml-et: Fix MUL_MAT MUL_MAT_ID remainders * ggml-et: Fix ET-SOC reference * ggml-et: Fix embed kernels scripts for old python This allows GGML-ET to build on pre-3.8 python. * Add sysemu support with compile time flag `-DGGML_ET_SYSEMU=ON` (#6) * Example using ET-Soc-1 emulator configuration Example usage: ```bash cmake -B build -DGGML_CUDA=OFF -DGGML_ET=ON -DLLAMA_CURL=OFF -DGGML_CCACHE=ON cmake --build build --config Release -j $(nproc) time ./build/bin/test-backend-ops ./build/bin/llama-server \ --model Qwen3-0.6B-Q8_0.gguf \ --alias Qwen3-0.6B-Q8_0 \ -fa 0 \ --ctx-size 1024 \ --no-warmup \ --host 127.0.0.1 \ --port 8080 ``` * build: proper dep tracking for kernels * support host using MOLD linker * initial multi core GET_ROW F32 implementation * vectorized q8 dequant * wip: cland warning clenaups and initial logging refactor * wip: message default message cleanup * chore: message cleanups * cmake cleanup * migrate to use platform provided functions * cmake back into subdir * support et_print() in kernels * fix: repair kernel building * perf: operations run async by default * debug: proper kernel dep tracking and error detection on kenrel launch * fix: kernel binary dep tracking and fixing get_rows_f32 erroring * perf: back to doing async kernel runs by default * perf: vectorize and parallel device memset * merge matmul work * misc: align allocation and enable all offload * misc: delete deadcode and respect memory limits * fix: repair tensor debug print * fix: loosen RMS_NORM op percision * feat: Q4_0 GET_ROWS * perf: FP32 MUL_MAT using TensorFMA * update limitations * perf: redue L1 load in compute_block_dot_product_q8_0 * feat: save kernel mapping (name to id) when profiling is enabled * chore: memops cleanup * perf: parallelize softmax by rows * perf: vectorize 2nd phase of softmax * perf: ban GET_ROWS from offloaded * perf: vectorize and non-atomic for eltwise ops and sub support * perf: vectorize normal rope * perf: glu runs in parallel * merge: manually merge saqib's work on kernel fixes * perf: more vectorized RoPE * perf: parallelize mul_mat_id * perf: parallelize set_rows_f32 * perf: vectorize softmax * feat: support kernel fusion and fuse RMS_NORM + MUL * fix: mostly resolve test-backend-ops failure in SOFT_MAX and ROPE * fix: bump max rope dims for gemma * feat: GeGLU and SCALE support to fully offload Gemma * perf: faster device memset * feat: get_rows supporting Q4_K and avoid cont cache coherent issues * better F32 MM * feat: NORM for ET backend * feat: SQR for ET backend * feat: UNARY on ET * feat: el_map support broadcasting for ET * feat: SUM_ROWS in ET backend * feat: more ops in ET backend * feat: WKV* operators in ET backend * perf: parallelize operators across cacheline instead of row * perf: parallelize get_rows on cacheline * wip: baseline FlashAttention for ET backend * wip: enough FA and CPY f32->f16 to run llama 3.1 fully offloaded with FA on * feat: f16 x f16 -> f32 MM using matrix engine * wip: f16 FlashAttention using matrix engine * wip: clean up * feat: barriers * perf: optimize FA_F16 in ET * perf: vectorize pack_k_for_transpose16 * perf: prefetch next loop matrix tile * perf: FlashAttention 2nd MM uses TensorFMA and optimizations * cleanup: flashattention reorg * perf: optimizations and fixes * feat: L2SCP API and make FlashAttention support DV = 256 for gemma * perf: parallelize norms beyond single row * feat: GATED_DELTA_NET support and relaxed L2_NORM requirment * feat: loosen RMS_NORM, NORM, ROPE contingous req too * feat: repeat supports brocasting on dim 0 and loosen cont check * feat: FILL and DIAG operator * feat: loosen UNARY support chcek * feat: TRI support * feat: SOLVE_TRI support * feat: basic SET support * feat: loosen CONT req * perf: fp16_to_fp32 use ASM * feat: IMROPE support * feat: PAD support * feat: global barrier * fix: view must live on the same backend as backing tensor * feat: relax CONCAT in ET backend * feat: dead simple CUMSUM implementation * feat: basic SSM_CONV support * feat: loosen CONCAT req * feat: relax GATED_DELTA_NET and add SET support proper * cleanup: cleanup LCM math * feat: SWIGLU single input * feat: SSM_SCAN support * feat: el_map supports non aligned tensors in best effort * feat: basic GROUP_NORM support * feat: loosen MUL_MAT capablities slightly * feat: loosen MUL_MAT and GET_ROWS and add IM2COL * feat: special case for softmax 1x1x1x1 * feat: loosen SOFT_MAX req in ET backend * fix: el_map unaligned acse fixes * perf: optimize zero_acc_vec in flash_attn_ext_f16_me * perf: use hart 1 for packing in MM and FA for FP16 * feat: kernel semaphore * perf: better instruction sequence in FlashAttention * fix: gated_delta_net with proper masking * perf: better parallelization for GATED_DELTA_NET * perf: parallelize SSM_CONV over nr * perf: vectorize SSM_CONV * perf: optimize MUL_MAT for q8 * feat: support Gemma 4 * fix: support multi-device * feat: broader GLU support * feat: unary ops supports view * fix: repair fp16 MM using matrix engine * perf: handle large N GEMV better * perf: better q8_0 MM * perf: better set_rows * add back deleted files * fix: repair after merge * feat: POC version of uberkernel * feat: RMS_NORM in uberkernel * feat: add more kernels into usage * chore: clean up uberkernel compilation * perf: faster flash attention * perf: opt flash attention for large seq length * feat: loosen op bounds. clamp and mean support * perf: vectorize ssm_scan * perf: slightly faster FA * perf: FlashAttention parallel MM and load * perf: fuse Q8 MM and ADD * feat: basic conv kernel for ET * softMAx_test * set_rows_f32 * get_rows and cont * testing * set_rows_exp * Junk addition * Narrowing the issue * Update flash_attn_ext_f16_me.c Focusing FA_ext_f16_me * test * Eviction updated * Detailed cache eviction debug * mulmat * removeal of `BUILD_FOR_UBERKERNEL` flag * cleaning... * fix: balance FCC0 count * feat: implement mul_mat and mul_mat_id for Q4_0 type * optimize uberkernel plan upload * add mul_mat q4 into uberkernel * enable gating flush to just uberkernel * update docs for ET * update op support for ET * et-backend: optimize Q4_0 and Q8_0 mul_mat_id row accumulations * et-backend: specialize mul_mat_id kernels for Q4_0 and Q8_0 * et-backend: fix RoPE YaRN corr_dim formula and handle degenerate inputs * test-backend-ops: add DeepSeek-V2-Lite RoPE test coverage * et-backend: add Q4_0 mul_mat matrix-engine kernel using TensorFMA32 * et-backend: vectorize Q4_0 matrix-engine dequantization * et-backend: support hybrid matrix/vector engine execution for Q4_0 mul_mat tail * et-backend: run partial-N tiles on matrix engine for Q4_0 mul_mat * et-backend: route Q4_0 mul_mat N < 53 to vecdot for better prefill latency * Update uberkernel.c * Update unary_f32.c * gemma 4 * bisect gemma4: enable scale_f32 only * bisect gemma4: +rms_norm_f32 * bisect gemma4: +rms_norm_mul_f32 * bisect gemma4: disable rms_norm_mul_f32 -- BREAKS OUTPUT * bisect gemma4: +rope_f32 (skip rms_norm_mul) * bisect gemma4: +el_map_f32 * bisect gemma4: +softmax_f32 * bisect gemma4: +get_rows_f32 * bisect gemma4: +glu_f32 * bisect gemma4: +mul_mat_f32 +mul_mat_f32_matrix_engine * bisect gemma4: +mul_mat_f16 +mul_mat_f16_matrix_engine * bisect gemma4: +mul_mat_Q8_0 +mul_mat_Q4_0 * bisect gemma4: +flash_attn_ext_f32 +flash_attn_ext_f16_me * bisect gemma4: +mul_mat_id_f32 * bisect gemma4: +sum_rows_f32 * bisect gemma4: +cont_f16 * bisect gemma4: +fill_f32 * bisect gemma4: +unary_f32 (all ops re-enabled except rms_norm_mul) * Update rms_norm_mul_f32.c * bisect2 gemma4 n64: +scale_f32 only * bisect2 gemma4 n64: +rms_norm_f32 +rope_f32 * bisect2 gemma4 n64: +rms_norm_mul_f32 (with ET_UBERKERNEL eviction fix) * bisect2 gemma4 n64: +el_map +get_rows +glu +softmax (skip rms_norm_mul) * bisect2 gemma4 n64: all ops enabled except rms_norm_mul * bisect2 n64: test unary+cont+fill+sum_rows (no mul_mat/flash_attn) * bisect2 n64: +mul_mat_f32 +mul_mat_f32_matrix_engine * bisect2 n64: +mul_mat_f16 +mul_mat_f16_matrix_engine * bisect2 n64: +mul_mat_Q8_0 +mul_mat_Q4_0 * bisect2 n64: +mul_mat_Q8_0 only (disable Q4_0) * bisect2 n64: +mul_mat_Q4_0 only (Q8_0 breaks) * bisect2 n64: +mul_mat_id +flash_attn_ext (skip Q8_0) * run-3: matmul + rms_norm_mul * run-4 * Revert "run-4" * run5 * changes after cleanup * cleanup before upstream * restrict changes into ET backend * move kernel embedding from Python to CMake * move uberkernel gen into CMake * apply clang format * update CMake style * update to match C and C++ style * use source ggml and quant headers instead of ET's * MROPE support * absorb view ops into same branch as none * fix bad rebase * add marty1885 to codeowners * oops * remove redundant newline * fix CI editor warnings --------- Co-authored-by: Vidas <vidas@nuolat.lt> Co-authored-by: Gianluca Guida <glguida@tlbflush.org> Co-authored-by: Gianluca Guida <gianluca@nekko.ai> Co-authored-by: ubergarm <leimgrub@gmail.com> Co-authored-by: SaqibAkram-10xE <saqib.akram@10xengineers.ai> Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
509 lines
20 KiB
C++
509 lines
20 KiB
C++
#include "ggml-et-kernels.h"
|
|
|
|
#include "ggml-et-kernels-embed.hpp"
|
|
#include "ggml-et-uberkernel-kernel-map.h"
|
|
#include "ggml-impl.h"
|
|
|
|
#include <cstdlib>
|
|
#include <cstring>
|
|
#include <fstream>
|
|
|
|
#define ET_TRACE_DECODER_IMPL
|
|
#include <et-trace/decoder.h>
|
|
#include <et-trace/layout.h>
|
|
|
|
static constexpr size_t GGML_ET_UBERKERNEL_PARAM_ALIGN = 64;
|
|
|
|
static size_t ggml_et_align_up(size_t value, size_t alignment) {
|
|
return (value + alignment - 1) & ~(alignment - 1);
|
|
}
|
|
|
|
static size_t ggml_et_next_capacity(size_t current_capacity, size_t required_capacity) {
|
|
if (current_capacity == 0) {
|
|
return required_capacity;
|
|
}
|
|
|
|
size_t next_capacity = current_capacity;
|
|
while (next_capacity < required_capacity) {
|
|
next_capacity *= 2;
|
|
}
|
|
|
|
return next_capacity;
|
|
}
|
|
|
|
static ggml_backend_et_uberkernel_slot & ggml_et_uberkernel_current_slot(ggml_backend_et_uberkernel_context * uk_ctx) {
|
|
return uk_ctx->slots[uk_ctx->current_slot];
|
|
}
|
|
|
|
// Wait for any in-flight launch that previously used this slot to finish,
|
|
// so the host vectors and device buffers are safe to mutate / free.
|
|
static void ggml_et_uberkernel_slot_wait(ggml_backend_et_uberkernel_slot & slot,
|
|
const std::shared_ptr<rt::IRuntime> & runtime) {
|
|
if (!slot.has_pending || !runtime) {
|
|
return;
|
|
}
|
|
runtime->waitForEvent(slot.pending_event);
|
|
slot.has_pending = false;
|
|
}
|
|
|
|
static void ggml_et_uberkernel_reset_segment(ggml_backend_et_uberkernel_context * uk_ctx) {
|
|
if (!uk_ctx) {
|
|
return;
|
|
}
|
|
|
|
uk_ctx->shire_mask = 0;
|
|
auto & slot = ggml_et_uberkernel_current_slot(uk_ctx);
|
|
// Drain any prior launch on this slot before clearing its host buffers.
|
|
// begin_graph and abort_graph both come through here; in either case we
|
|
// must not yank the source memory out from under an in-flight DMA.
|
|
ggml_et_uberkernel_slot_wait(slot, ggml_et_runtime());
|
|
slot.insts.clear();
|
|
slot.params_blob.clear();
|
|
}
|
|
|
|
static bool ggml_et_uberkernel_ensure_slot_capacity(ggml_backend_et_uberkernel_slot & slot,
|
|
ggml_backend_et_device_context * dev_ctx,
|
|
size_t insts_size,
|
|
size_t params_size) {
|
|
std::shared_ptr<rt::IRuntime> runtime = ggml_et_runtime();
|
|
if (!dev_ctx || !runtime) {
|
|
return false;
|
|
}
|
|
|
|
try {
|
|
if (slot.device_insts == nullptr || insts_size > slot.device_insts_capacity) {
|
|
const size_t new_capacity = ggml_et_next_capacity(slot.device_insts_capacity, insts_size);
|
|
if (slot.device_insts) {
|
|
runtime->freeDevice(dev_ctx->rtid, slot.device_insts);
|
|
}
|
|
slot.device_insts = runtime->mallocDevice(dev_ctx->rtid, new_capacity);
|
|
slot.device_insts_capacity = slot.device_insts ? new_capacity : 0;
|
|
}
|
|
|
|
if (slot.device_params == nullptr || params_size > slot.device_params_capacity) {
|
|
const size_t new_capacity = ggml_et_next_capacity(slot.device_params_capacity, params_size);
|
|
if (slot.device_params) {
|
|
runtime->freeDevice(dev_ctx->rtid, slot.device_params);
|
|
}
|
|
slot.device_params = runtime->mallocDevice(dev_ctx->rtid, new_capacity);
|
|
slot.device_params_capacity = slot.device_params ? new_capacity : 0;
|
|
}
|
|
} catch (const std::exception & e) {
|
|
GGML_LOG_ERROR("ET: Failed to resize uberkernel buffers: %s\n", e.what());
|
|
return false;
|
|
}
|
|
|
|
return slot.device_insts != nullptr && slot.device_params != nullptr;
|
|
}
|
|
|
|
// Get embedded kernel data by name
|
|
static std::vector<std::byte> ggml_et_get_embedded_kernel(const std::string & kernel_name) {
|
|
auto it = ggml_et_embedded_kernels.find(kernel_name);
|
|
if (it == ggml_et_embedded_kernels.end()) {
|
|
GGML_LOG_ERROR("ET: Unknown embedded kernel: %s\n", kernel_name.c_str());
|
|
return {};
|
|
}
|
|
|
|
const unsigned char * data = it->second.first;
|
|
uint64_t size = it->second.second;
|
|
|
|
std::vector<std::byte> buffer(size);
|
|
std::memcpy(buffer.data(), data, size);
|
|
|
|
return buffer;
|
|
}
|
|
|
|
// Read kernel from file (for development/override)
|
|
static std::vector<std::byte> ggml_et_read_kernel_file(const std::string & kernel_path) {
|
|
std::ifstream file(kernel_path, std::ios::binary | std::ios::ate);
|
|
if (!file) {
|
|
return {};
|
|
}
|
|
|
|
auto size = file.tellg();
|
|
file.seekg(0, std::ios::beg);
|
|
|
|
std::vector<std::byte> buffer(size);
|
|
file.read(reinterpret_cast<char *>(buffer.data()), size);
|
|
|
|
return buffer;
|
|
}
|
|
|
|
// Load kernel from file or embedded data
|
|
bool ggml_et_load_kernel(ggml_backend_et_device_context * dev_ctx, const std::string & kernel_name) {
|
|
std::shared_ptr<rt::IRuntime> runtime = ggml_et_runtime();
|
|
if (!runtime) {
|
|
GGML_LOG_ERROR("ET: Runtime not available for kernel loading\n");
|
|
return false;
|
|
}
|
|
|
|
// Check if kernel already loaded
|
|
if (dev_ctx->loaded_kernels.find(kernel_name) != dev_ctx->loaded_kernels.end()) {
|
|
GGML_LOG_DEBUG("ET: Kernel %s already loaded on device %d\n", kernel_name.c_str(), dev_ctx->devidx);
|
|
return true;
|
|
}
|
|
|
|
std::vector<std::byte> kernel_data;
|
|
const char * kernels_path = getenv("GGML_ET_KERNELS_PATH");
|
|
|
|
// If GGML_ET_KERNELS_PATH is set, try to load from file first
|
|
if (kernels_path) {
|
|
std::string kernel_file = std::string(kernels_path) + "/" + kernel_name + ".elf";
|
|
kernel_data = ggml_et_read_kernel_file(kernel_file);
|
|
|
|
if (!kernel_data.empty()) {
|
|
GGML_LOG_INFO("ET: Loading kernel %s from file: %s\n", kernel_name.c_str(), kernel_file.c_str());
|
|
} else {
|
|
GGML_LOG_INFO("ET: Kernel file not found: %s, falling back to embedded\n", kernel_file.c_str());
|
|
}
|
|
}
|
|
|
|
// If no file data, use embedded kernel
|
|
if (kernel_data.empty()) {
|
|
kernel_data = ggml_et_get_embedded_kernel(kernel_name);
|
|
if (kernel_data.empty()) {
|
|
GGML_LOG_ERROR("ET: Failed to get kernel data for %s\n", kernel_name.c_str());
|
|
return false;
|
|
}
|
|
}
|
|
|
|
try {
|
|
// Load kernel code using device's default stream
|
|
auto load_result = runtime->loadCode(dev_ctx->default_stream, kernel_data.data(), kernel_data.size());
|
|
runtime->waitForEvent(load_result.event_);
|
|
|
|
// Store kernel handle
|
|
dev_ctx->loaded_kernels[kernel_name] = load_result.kernel_;
|
|
return true;
|
|
|
|
} catch (const std::exception & e) {
|
|
GGML_LOG_ERROR("ET: Failed to load kernel %s: %s\n", kernel_name.c_str(), e.what());
|
|
return false;
|
|
}
|
|
}
|
|
|
|
static bool ggml_et_launch_kernel_internal(ggml_backend_et_device_context * dev_ctx,
|
|
const std::string & kernel_name,
|
|
void * params,
|
|
size_t params_size,
|
|
uint64_t shire_mask,
|
|
bool enable_print,
|
|
bool sync_error_check,
|
|
rt::EventId * out_event = nullptr) {
|
|
std::shared_ptr<rt::IRuntime> runtime = ggml_et_runtime();
|
|
if (!runtime) {
|
|
GGML_LOG_ERROR("ET: Runtime not available for kernel launch\n");
|
|
return false;
|
|
}
|
|
|
|
// Lazy loading: check if kernel is loaded, load if needed
|
|
auto kernel_it = dev_ctx->loaded_kernels.find(kernel_name);
|
|
if (kernel_it == dev_ctx->loaded_kernels.end()) {
|
|
// Kernel not loaded - load it
|
|
if (!ggml_et_load_kernel(dev_ctx, kernel_name)) {
|
|
GGML_LOG_ERROR("ET: Failed to lazy-load kernel %s\n", kernel_name.c_str());
|
|
return false;
|
|
}
|
|
|
|
// Update iterator after successful load
|
|
kernel_it = dev_ctx->loaded_kernels.find(kernel_name);
|
|
if (kernel_it == dev_ctx->loaded_kernels.end()) {
|
|
GGML_LOG_ERROR("ET: Kernel %s not found after loading\n", kernel_name.c_str());
|
|
return false;
|
|
}
|
|
}
|
|
|
|
rt::KernelId kernel_id = kernel_it->second;
|
|
|
|
try {
|
|
// Setup kernel launch options
|
|
rt::KernelLaunchOptions k_opts;
|
|
k_opts.setShireMask(shire_mask); // Default: all shires (0xFFFFFFFF)
|
|
k_opts.setBarrier(true); // Wait for completion
|
|
k_opts.setFlushL3(false); // No L3 flush needed
|
|
if (enable_print) {
|
|
k_opts.setUserTracing(reinterpret_cast<uint64_t>(dev_ctx->trace_buffer),
|
|
static_cast<uint32_t>(ET_TRACE_BUFFER_SIZE),
|
|
0, // threshold
|
|
shire_mask, // shire mask
|
|
0xFFFFFFFFFFFFFFFFULL, // threadMask - all threads
|
|
0xFFFFFFFFU, // eventMask - all events
|
|
0xFFFFFFFFU // filterMask - all levels
|
|
);
|
|
}
|
|
|
|
if (sync_error_check) {
|
|
runtime->waitForStream(dev_ctx->default_stream);
|
|
auto errors = runtime->retrieveStreamErrors(dev_ctx->default_stream);
|
|
if (!errors.empty()) {
|
|
GGML_LOG_ERROR("ET: Errors detected before kernel \"%s\" launch\n", kernel_name.c_str());
|
|
for (const auto & error : errors) {
|
|
GGML_LOG_ERROR("ET: Error code: %d\n", (int) error.errorCode_);
|
|
}
|
|
abort();
|
|
}
|
|
}
|
|
|
|
rt::EventId launch_event = runtime->kernelLaunch(dev_ctx->default_stream, kernel_id,
|
|
reinterpret_cast<std::byte *>(params), params_size, k_opts);
|
|
if (out_event) {
|
|
*out_event = launch_event;
|
|
}
|
|
|
|
if (enable_print) {
|
|
std::vector<std::byte> host_trace_buf(ET_TRACE_BUFFER_SIZE);
|
|
runtime->memcpyDeviceToHost(dev_ctx->default_stream, dev_ctx->trace_buffer, host_trace_buf.data(),
|
|
ET_TRACE_BUFFER_SIZE);
|
|
runtime->waitForStream(dev_ctx->default_stream);
|
|
const auto * trace_header = reinterpret_cast<const trace_buffer_std_header_t *>(host_trace_buf.data());
|
|
const trace_entry_header_t * entry = nullptr;
|
|
while ((entry = Trace_Decode(trace_header, entry))) {
|
|
if (entry->type != TRACE_TYPE_STRING) {
|
|
continue;
|
|
}
|
|
const auto * str_entry = reinterpret_cast<const trace_string_t *>(entry);
|
|
printf("[hart %d] %s", entry->hart_id, str_entry->string);
|
|
}
|
|
}
|
|
|
|
if (sync_error_check) {
|
|
// Already triggered. No need to retrigger
|
|
if (!enable_print) {
|
|
runtime->waitForStream(dev_ctx->default_stream);
|
|
}
|
|
auto errors = runtime->retrieveStreamErrors(dev_ctx->default_stream);
|
|
if (!errors.empty()) {
|
|
GGML_LOG_ERROR("ET: Errors detected during kernel \"%s\" execution\n", kernel_name.c_str());
|
|
for (const auto & error : errors) {
|
|
GGML_LOG_ERROR("ET: Error code: %d\n", (int) error.errorCode_);
|
|
}
|
|
abort();
|
|
}
|
|
}
|
|
|
|
return true;
|
|
} catch (const std::exception & e) {
|
|
GGML_LOG_ERROR("ET: Failed to launch kernel %s: %s\n", kernel_name.c_str(), e.what());
|
|
return false;
|
|
}
|
|
}
|
|
|
|
void ggml_et_uberkernel_begin_graph(ggml_backend_et_uberkernel_context * uk_ctx) {
|
|
if (!uk_ctx) {
|
|
return;
|
|
}
|
|
|
|
uk_ctx->failed = false;
|
|
ggml_et_uberkernel_reset_segment(uk_ctx);
|
|
}
|
|
|
|
static bool ggml_et_launch_uberkernel_segment(ggml_backend_et_device_context * dev_ctx,
|
|
ggml_backend_et_uberkernel_context * uk_ctx) {
|
|
if (!uk_ctx || !dev_ctx) {
|
|
return false;
|
|
}
|
|
|
|
auto & slot = ggml_et_uberkernel_current_slot(uk_ctx);
|
|
if (slot.insts.empty()) {
|
|
return true;
|
|
}
|
|
|
|
std::shared_ptr<rt::IRuntime> runtime = ggml_et_runtime();
|
|
if (!runtime) {
|
|
GGML_LOG_ERROR("ET: Runtime not available for uberkernel commit\n");
|
|
uk_ctx->failed = true;
|
|
return false;
|
|
}
|
|
|
|
const size_t insts_size = slot.insts.size() * sizeof(ggml_et_uberkernel_inst);
|
|
const size_t params_size = slot.params_blob.size();
|
|
const uint64_t shire_mask = uk_ctx->shire_mask;
|
|
bool ok = false;
|
|
|
|
try {
|
|
if (!ggml_et_uberkernel_ensure_slot_capacity(slot, dev_ctx, insts_size, params_size)) {
|
|
GGML_LOG_ERROR("ET: Failed to allocate uberkernel device buffers\n");
|
|
uk_ctx->failed = true;
|
|
// Drop this segment but keep the slot drained so we don't leak
|
|
// host vectors into the next graph.
|
|
slot.insts.clear();
|
|
slot.params_blob.clear();
|
|
uk_ctx->shire_mask = 0;
|
|
return false;
|
|
}
|
|
|
|
// Fire-and-forget H2D + launch on default_stream. In-stream FIFO
|
|
// ordering guarantees the kernel sees fully-uploaded buffers; the
|
|
// host source bytes (slot.insts / slot.params_blob) stay alive
|
|
// because we won't touch this slot again until pending_event fires.
|
|
runtime->memcpyHostToDevice(dev_ctx->default_stream, reinterpret_cast<const std::byte *>(slot.insts.data()),
|
|
slot.device_insts, insts_size, true);
|
|
runtime->memcpyHostToDevice(dev_ctx->default_stream, slot.params_blob.data(), slot.device_params, params_size,
|
|
true);
|
|
|
|
ggml_et_uberkernel_params params = {
|
|
static_cast<uint32_t>(slot.insts.size()),
|
|
static_cast<uint32_t>(sizeof(ggml_et_uberkernel_inst)),
|
|
reinterpret_cast<uint64_t>(slot.device_insts),
|
|
reinterpret_cast<uint64_t>(slot.device_params),
|
|
};
|
|
|
|
rt::EventId launch_event{};
|
|
ok = ggml_et_launch_kernel_internal(dev_ctx, "uberkernel", ¶ms, sizeof(params), shire_mask, false, false,
|
|
&launch_event);
|
|
if (ok) {
|
|
// The kernelLaunch above is the last thing on default_stream
|
|
// that touches this slot's device buffers. Recording its event
|
|
// lets the next reuse of this slot wait on that one event
|
|
// instead of the whole stream.
|
|
slot.pending_event = launch_event;
|
|
slot.has_pending = true;
|
|
}
|
|
} catch (const std::exception & e) {
|
|
GGML_LOG_ERROR("ET: Failed to commit uberkernel segment: %s\n", e.what());
|
|
}
|
|
uk_ctx->failed = !ok;
|
|
|
|
if (ok) {
|
|
uk_ctx->current_slot = (uk_ctx->current_slot + 1) % ggml_backend_et_uberkernel_context::SLOT_COUNT;
|
|
auto & next = ggml_et_uberkernel_current_slot(uk_ctx);
|
|
ggml_et_uberkernel_slot_wait(next, runtime);
|
|
next.insts.clear();
|
|
next.params_blob.clear();
|
|
} else {
|
|
slot.insts.clear();
|
|
slot.params_blob.clear();
|
|
}
|
|
uk_ctx->shire_mask = 0;
|
|
return ok;
|
|
}
|
|
|
|
void ggml_et_uberkernel_abort_graph(ggml_backend_et_uberkernel_context * uk_ctx) {
|
|
if (!uk_ctx) {
|
|
return;
|
|
}
|
|
|
|
uk_ctx->failed = false;
|
|
ggml_et_uberkernel_reset_segment(uk_ctx);
|
|
}
|
|
|
|
bool ggml_et_uberkernel_failed(const ggml_backend_et_uberkernel_context * uk_ctx) {
|
|
return uk_ctx && uk_ctx->failed;
|
|
}
|
|
|
|
static bool ggml_et_launch_uberkernel(ggml_backend_et_device_context * dev_ctx,
|
|
const std::string & kernel_name,
|
|
void * params,
|
|
size_t params_size,
|
|
uint64_t shire_mask,
|
|
bool enable_print,
|
|
bool sync_error_check) {
|
|
if (!dev_ctx) {
|
|
return false;
|
|
}
|
|
|
|
ggml_backend_et_uberkernel_context * uk_ctx = &dev_ctx->uberkernel;
|
|
const uint16_t uberkernel_id = ggml_et_uberkernel_kernel_id_from_name(kernel_name.c_str());
|
|
if (uberkernel_id == GGML_ET_UBERKERNEL_KERNEL_INVALID) {
|
|
if (!ggml_et_launch_uberkernel_segment(dev_ctx, uk_ctx)) {
|
|
return false;
|
|
}
|
|
return ggml_et_launch_kernel_internal(dev_ctx, kernel_name, params, params_size, shire_mask, enable_print,
|
|
sync_error_check);
|
|
}
|
|
|
|
auto & slot = ggml_et_uberkernel_current_slot(uk_ctx);
|
|
const size_t params_offset = ggml_et_align_up(slot.params_blob.size(), GGML_ET_UBERKERNEL_PARAM_ALIGN);
|
|
if (params_offset > slot.params_blob.size()) {
|
|
slot.params_blob.resize(params_offset);
|
|
}
|
|
|
|
const std::byte * params_bytes = reinterpret_cast<const std::byte *>(params);
|
|
slot.params_blob.insert(slot.params_blob.end(), params_bytes, params_bytes + params_size);
|
|
|
|
ggml_et_uberkernel_inst inst = {
|
|
uberkernel_id,
|
|
0,
|
|
static_cast<uint32_t>(params_offset),
|
|
static_cast<uint32_t>(params_size),
|
|
};
|
|
slot.insts.push_back(inst);
|
|
|
|
if (slot.insts.size() == 1) {
|
|
uk_ctx->shire_mask = shire_mask;
|
|
}
|
|
|
|
return true;
|
|
}
|
|
|
|
bool ggml_et_uberkernel_end_graph(ggml_backend_et_device_context * dev_ctx) {
|
|
if (!dev_ctx || !dev_ctx->uberkernel_enabled) {
|
|
return true;
|
|
}
|
|
|
|
return ggml_et_launch_uberkernel_segment(dev_ctx, &dev_ctx->uberkernel);
|
|
}
|
|
|
|
bool ggml_et_launch_kernel(ggml_backend_et_device_context * dev_ctx,
|
|
const std::string & kernel_name,
|
|
void * params,
|
|
size_t params_size,
|
|
uint64_t shire_mask,
|
|
bool enable_print,
|
|
bool sync_error_check) {
|
|
if (!dev_ctx) {
|
|
return false;
|
|
}
|
|
|
|
if (!dev_ctx->uberkernel_enabled) {
|
|
return ggml_et_launch_kernel_internal(dev_ctx, kernel_name, params, params_size, shire_mask, enable_print,
|
|
sync_error_check);
|
|
}
|
|
|
|
return ggml_et_launch_uberkernel(dev_ctx, kernel_name, params, params_size, shire_mask, enable_print,
|
|
sync_error_check);
|
|
}
|
|
|
|
void ggml_et_unload_kernel(ggml_backend_et_device_context * dev_ctx, const std::string & kernel_name) {
|
|
std::shared_ptr<rt::IRuntime> runtime = ggml_et_runtime();
|
|
if (!runtime) {
|
|
return;
|
|
}
|
|
|
|
auto kernel_it = dev_ctx->loaded_kernels.find(kernel_name);
|
|
if (kernel_it != dev_ctx->loaded_kernels.end()) {
|
|
try {
|
|
runtime->unloadCode(kernel_it->second);
|
|
dev_ctx->loaded_kernels.erase(kernel_it);
|
|
} catch (const std::exception & e) {
|
|
GGML_LOG_ERROR("ET: Failed to unload kernel %s: %s\n", kernel_name.c_str(), e.what());
|
|
}
|
|
}
|
|
}
|
|
|
|
void ggml_et_unload_all_kernels(ggml_backend_et_device_context * dev_ctx) {
|
|
if (!dev_ctx) {
|
|
return;
|
|
}
|
|
|
|
// Make a copy of kernel names since ggml_et_unload_kernel modifies the map
|
|
std::vector<std::string> kernel_names;
|
|
kernel_names.reserve(dev_ctx->loaded_kernels.size());
|
|
for (const auto & kernel_pair : dev_ctx->loaded_kernels) {
|
|
kernel_names.push_back(kernel_pair.first);
|
|
}
|
|
|
|
for (const auto & kernel_name : kernel_names) {
|
|
ggml_et_unload_kernel(dev_ctx, kernel_name);
|
|
}
|
|
}
|
|
|
|
std::vector<std::pair<std::string, rt::KernelId>> ggml_et_get_loaded_kernels(ggml_backend_et_device_context * dev_ctx) {
|
|
std::vector<std::pair<std::string, rt::KernelId>> loaded_kernels;
|
|
loaded_kernels.reserve(dev_ctx->loaded_kernels.size());
|
|
for (const auto & kernel_pair : dev_ctx->loaded_kernels) {
|
|
loaded_kernels.push_back(kernel_pair);
|
|
}
|
|
return loaded_kernels;
|
|
}
|