WEBINAR: Porting, Analysing & Optimising AI Models onto Edge GPUs

Running AI models efficiently at the edge is about more than hardware TOPS. In this session, we’ll walk through the end-to-end process of porting, analysing and optimising modern AI models for edge GPUs, using Qwen 3.5 as a practical example. Learn how to take a model from framework to deployment, identify performance bottlenecks, leverage optimised kernel libraries, and use profiling tools to maximise throughput, responsiveness and efficiency on edge devices. We’ll also explore how a programmable GPU architecture helps developers adapt as models, operators and AI workloads continue to evolve.
Presented by Mario Florea. Hosted by SemiWiki.
Mario Florea is Director of Engineering at Imagination Technologies, where he leads engineering efforts focused on AI software for GPU-accelerated edge inference. His work centers on enabling efficient, high-performance AI workloads on resource-constrained devices, bridging the gap between modern AI models and highly optimized GPU hardware.
With a strong background in engineering and technology leadership, Mario works at the intersection of AI software, and GPU architecture, helping turn emerging AI capabilities into practical solutions for the edge.
Share this post via:









ASML Has High-NA and Chipmakers Can’t Say No