S&P 500 5,235.18 +1.02%EUR/USD 1.0840 +0.21%GBP/USD 1.2710 +0.14%USD/JPY 149.50 −0.18%BRENT $82.40 −0.81%BTC $67,800 −0.21%GOLD $2,341 +0.55%NASDAQ 16,420.55 +0.74%S&P 500 5,235.18 +1.02%EUR/USD 1.0840 +0.21%GBP/USD 1.2710 +0.14%USD/JPY 149.50 −0.18%BRENT $82.40 −0.81%BTC $67,800 −0.21%GOLD $2,341 +0.55%NASDAQ 16,420.55 +0.74%
A daily business newspaper · Founded in 2026

Money Talk

Finance and markets: business, quotes, gold, energy and releases.

Moreh Demonstrates High-Speed LLM Inference on AMD Instinct GPUs

At the AMD Advancing AI 2026 event in San Francisco, software developer Moreh showcased its MoAI Inference Framework running the GLM-5.1 model on 32 AMD Instinct MI300X GPUs. By displaying real-time metrics like tokens per second and GPU utilization, the company aimed to prove production-grade efficiency on AMD hardware.

Moreh Demonstrates High-Speed LLM Inference on AMD Instinct GPUs
Photo: Bio & News

The demonstration utilized four nodes to run the computationally heavy GLM-5.1 model, allowing attendees to interact with the chatbot while monitoring performance data. Unlike static showcases, this setup provided transparency regarding Time To First Token (TTFT) and Time Per Output Token (TPOT), metrics critical for enterprise-scale AI deployment.

Moreh’s framework is currently the only commercial distributed inference solution tailored for the AMD ecosystem. CEO Gangwon Jo noted that the company intends to lower infrastructure costs by optimizing heterogeneous computing, effectively decoupling high-performance AI services from reliance on specific, monolithic hardware environments. Through its subsidiary, Motif Technologies, the firm is expanding its end-to-end capabilities, signaling a push to capture market share alongside partners like Tenstorrent.

Share article
TelegramXFacebook

When reusing this material a link to Money Talk is required.

Comments (0)

Leave a comment

No comments yet. Be the first!