Disclosure: This article began as a conversation with an OpenAI model. It has been edited for clarity and publication, and its time-sensitive technical claims were checked against current vendor and project documentation.

I was watching a video comparing AMD's Ryzen AI Halo platform with NVIDIA's DGX Spark when an obvious question came to mind: the AMD machine does not run CUDA, so does that mean it can run only open-source software specifically written for AMD GPUs?

The short answer is no. “Open source” is not the important boundary. The important question is whether the software has a computational backend compatible with the hardware.

The model is not the runtime

It helps to separate a local AI system into three layers:

  1. Model weights: Llama, Qwen, Mistral, Stable Diffusion and similar models.
  2. The application or runtime: PyTorch, llama.cpp, vLLM, Ollama, ComfyUI and related software.
  3. The hardware backend: CUDA, ROCm/HIP, Vulkan, DirectML, Metal or CPU execution.

The model weights are generally hardware-neutral. A GGUF model, for example, is not inherently an NVIDIA model or an AMD model. The program loading those weights must know how to send the computations to the available processor.

On an NVIDIA machine, that backend is commonly CUDA. On AMD hardware, it may be ROCm/HIP or Vulkan. Windows applications may also use DirectML. The same application can support several backends and choose the appropriate one at build time or runtime.

The llama.cpp project is a useful example. It supports NVIDIA GPUs through CUDA and AMD GPUs through HIP, while also offering Vulkan and other backends. The model can remain the same while the execution path changes. See the project's supported backends and build documentation.

What AMD uses instead of CUDA

CUDA is NVIDIA's GPU computing platform. An AMD GPU does not natively execute a CUDA binary.

AMD's principal compute stack is ROCm, and HIP is its CUDA-like C++ runtime and kernel language. AMD deliberately designed HIP to make CUDA projects easier to port. Its HIPIFY tools can translate many CUDA API calls into HIP equivalents.

That is still a source-code conversion and rebuild, not a magic compatibility switch. If a project ships only a precompiled CUDA extension, the AMD machine ordinarily cannot use it. If the source is available but deeply tied to NVIDIA-specific libraries or custom kernels, porting may range from straightforward to impractical. AMD explains this distinction in its HIP porting guide.

This produces four useful cases:

So “open source” helps because a project can be ported, but it does not itself provide AMD compatibility.

Why Ryzen AI Halo is interesting

AMD's Ryzen AI Halo developer platform combines a Ryzen AI Max+ processor, an integrated RDNA 3.5 GPU and a large pool of unified memory. AMD advertises the current platform with up to 128 GB of LPDDR5X memory and full ROCm support on Windows and Linux. The current ROCm compatibility matrix includes Ryzen AI Max systems based on Strix Halo. See AMD's Ryzen AI Halo overview and ROCm compatibility matrix.

That large shared-memory pool is the attraction for local language models. A quantized model that would not fit into the dedicated memory of a conventional consumer GPU may fit into a unified-memory system. Capacity, however, is not the same thing as speed. Memory bandwidth, backend optimization, model architecture and quantization all affect real performance.

The “Ryzen AI” name also includes an NPU, but the NPU is not automatically where large local language models run. Most of today's desktop LLM tools use the GPU or CPU unless they explicitly support AMD's NPU software stack.

Halo also uses the familiar x86-64 software environment. That can make ordinary desktop and development software easier to accommodate than on an Arm-based system.

What DGX Spark offers

NVIDIA's DGX Spark combines a 20-core Arm processor with a Blackwell GPU and 128 GB of coherent unified memory. NVIDIA positions it as a compact AI development system and supplies DGX OS, CUDA, cuDNN, container support and access to its NGC software ecosystem. NVIDIA lists frameworks including PyTorch and TensorRT-LLM among its supported AI tools. See NVIDIA's DGX Spark hardware overview and system overview.

Spark's largest advantage is therefore not merely a number on a hardware specification sheet. It is access to the CUDA ecosystem.

Much AI research and development still arrives CUDA-first. New custom kernels, training recipes, quantization techniques and experimental repositories frequently work on NVIDIA before they work elsewhere. Some eventually receive ROCm implementations; others require community ports; some remain CUDA-only.

Spark is not universally compatible, either. Its CPU is Arm64, so host-side packages compiled only for x86-64 can present a separate problem even when their GPU code supports CUDA. NVIDIA's curated containers reduce that friction, but “it uses CUDA” does not guarantee that every old desktop AI package will run unchanged.

Which machine is the safer choice?

For widely used quantized local language models, both platforms can be viable. Tools built on llama.cpp or another runtime with a mature AMD backend can make excellent use of Halo's unified memory. In that setting, the specific model, quantization, context size and measured token rate matter more than the CUDA brand by itself.

Ryzen AI Halo is especially attractive when you value:

DGX Spark is safer when you value:

The practical purchasing question is therefore not “Does AMD run AI?” It plainly does. The better question is: Does the software I actually intend to use have a well-maintained AMD backend?

If the answer is yes, Halo can be compelling. If the plan involves trying arbitrary GitHub repositories, specialized training code and the latest CUDA kernels, Spark offers the less hazardous compatibility path.

The concise answer

AMD Ryzen AI Halo does not run CUDA natively. It runs AI workloads through alternatives such as ROCm/HIP, Vulkan, DirectML or the CPU. The same model weights may run on AMD and NVIDIA, but the software executing them needs the appropriate backend.

Open source makes adaptation possible; it does not make CUDA code automatically portable. NVIDIA's advantage remains the breadth and maturity of the CUDA ecosystem. AMD's opportunity is to offer large-memory, x86-based local AI systems whose supported workloads can be remarkably capable.

That is the distinction hidden by many simple hardware comparisons: model compatibility, software compatibility and performance are three different questions.

The practical way to choose is to begin with the work rather than the vendor. List the runtimes, model formats, training frameworks and extensions you expect to use. Check those exact tools against current support matrices and real benchmarks. If they are well supported by ROCm, HIP or Vulkan, Halo may be an unusually flexible high-memory x86 system. If they assume CUDA, Spark reduces integration risk.

The best machine is the one that turns your intended software into working results—not necessarily the one with the cleanest headline specification.

Frank Kurka
[email protected]
kurkalabs.dev