Case-Study: V-Nova - GPU Porting from OpenCL to Metal
Dec 15, 2023
Case study on moving a GPU application from OpenCL to Metal for our client V-Nova.
Read moreWhy Choose Us?
Most GPU work handed to us starts with the wrong assumption about where the time goes. We measure the workload honestly, then change the algorithm, the memory layout, or the kernel, whichever actually moves the number.
Founder-Led GPU Expertise
Balázs Keszthelyi built the first OpenCL benchmark adopted by major GPU vendors and architected the VC-6 codec. With decades of GPU-first innovation, TechnoLynx is led by one of the field’s most credible pioneers.
Algorithm Redesign for Speed
We don’t just move code to GPUs, we rethink the algorithm. From simulation engines to custom AI pipelines, we redesign logic to unlock real-world speedups that straight-up GPU porting can’t deliver.
Full-Stack Performance Tuning
GPU speed means nothing if the rest of your stack lags. We optimise across CPU, memory, and I/O to eliminate bottlenecks and ensure your system performs as a whole.
AI + GPU: Smarter, Faster Systems
We blend custom-coded logic with AI inference, optimised for GPU performance engineering. The result: intelligent systems that are fast, efficient, and ready for real-time deployment.
Cross-Platform GPU Porting
From CUDA to Metal, OpenCL to Vulkan: we make your code run fast on any GPU. We’ve helped clients unlock Apple silicon, AMD, and NVIDIA platforms with precision.
Visual Computing, Not Just Compute
We don’t just accelerate code, we visualise it. From GPU-accelerated simulations to 3D rendering and XR, we bring deep graphics expertise to projects that need both performance and visual clarity.
Want This as a Packaged Engagement?
When the goal is a measured cost-per-request or latency saving on a workload already running in production, the packaged way to buy this is the Inference Cost-Cut Pack: an Audit that ranks where the cost leaks, then an Optimisation Sprint that builds the high-confidence changes and hands back a harness that reproduces the numbers.
Where This Sits
This page is engineering delivered to a client's spec; LynxBenchAI is the open methodology behind how we hold those numbers to a standard in the first place, sustained throughput, per precision, under a declared and bounded optimisation budget. Building the benchmark harness yourself is LynxBenchAI territory; the packaged outcome for your workload is the Inference Cost-Cut Pack above.
TechnoLynx delivered the project on time and provided quality outputs that met the client's expectations. The team was proactive in providing ideas and suggestions, and they were careful at properly planning the tasks. The client also praised the team's expertise in GPU programming and AI.
TechnoLynx's skill in low-level software development was impressive. TechnoLynx was able to create four prototypes with common components and an interface for easy maintenance. The client was extremely happy with the solution's speed. Moreover, their communication was seamless and straightforward.
TechnoLynx's unique aspect is that they're able to transform complex theories into practicable and applicable results. TechnoLynx provides research reports and architecture planning documents. The team is able to transform complex theories into practicable and applicable results. TechnoLynx's project management is strong and delivers work on time without hardware issues, being responsive through virtual meetings.
I’m delighted with our collaboration with their team. Thanks to TechnoLynx's work, the client has been able to co-author two patents. They lead responsive project management to solve problems quickly. The team also praises their skilled and knowledgeable team.
We had high-efficiency meetings. TechnoLynx’s work resulted in a successful breakthrough, and their input improved the client’s app. Their flexible and organised project management cultivated a healthy collaboration experience. Ultimately, their professionalism and commitment were impressive.
TechnoLynx stands out through pioneer-level leadership and proprietary innovation. Our founder, Balázs Keszthelyi, architected the VC-6 codec and built the first OpenCL benchmark adopted by major GPU vendors. This deep-tech heritage allows us to solve optimization challenges that standard engineering firms cannot.
TechnoLynx provides cross-platform optimization for a wide range of hardware and frameworks, including:
We ensure high-performance execution for AI and computer vision pipelines regardless of the underlying architecture.
We achieve low-latency AI inference through three core techniques: 1. Model Quantisation: Reducing precision without losing accuracy. 2. Pruning: Removing redundant parameters. 3. Custom Pipelines: Utilizing TensorRT and ONNX for maximum throughput on specific GPU architectures.
Yes, TechnoLynx designs future-proof, vendor-agnostic solutions. We ensure your GPU software is portable across different vendors (e.g., migrating from NVIDIA to AMD) and operating systems, providing long-term flexibility and scalability.
Yes, our XR services include:
Yes, we tune pipelines for the entire AI lifecycle. This includes multi-GPU setups for scalable training and using optimized runtimes like TensorRT for efficient deployment and inference.
We architect distributed solutions that achieve near-linear scaling. By optimizing inter-GPU communication, we enable high throughput for large-scale simulations, massive rendering tasks, and complex AI models.
We treat memory as a primary resource. Using advanced profiling, we analyze memory access patterns to minimize latency and eliminate bottlenecks, which is critical for processing large datasets in real-time.
Yes. Our auditing services include kernel profiling, shader analysis, and end-to-end system stress testing. We provide clients with actionable data and recommendations to maximize their hardware investment. The measurement discipline behind those numbers, performance read as a property of the whole hardware-and-software stack, sustained under load, reported per precision, is the methodology we develop in the open at LynxBenchAI.
Most of the time, once a workload is profiled honestly. Micro-optimisations rarely beat changing the algorithm or the data layout, because the biggest GPU gains usually come from removing whole memory round-trips or restructuring the dependency graph rather than from squeezing a few percent out of a single kernel. We always profile first to identify which level of leverage actually applies. See how to profile GPU kernels and when algorithmic restructuring beats kernel tuning.
Explore our latest thought leadership on innovation, technology, and industry best practices.
Dec 15, 2023
Case study on moving a GPU application from OpenCL to Metal for our client V-Nova.
Read more
May 15, 2023
How TechnoLynx modelled AI inference performance across GPU architectures — delivering two tools (topology-level performance predictor and OpenCL GPU…
Read more
Jan 23, 2020
TechnoLynx used GPU acceleration to improve physics simulations for an SME, leveraging dedicated graphics cards, advanced algorithms, and real-time…
Read more