The demonstration utilized four nodes to run the computationally heavy GLM-5.1 model, allowing attendees to interact with the chatbot while monitoring performance data. Unlike static showcases, this setup provided transparency regarding Time To First Token (TTFT) and Time Per Output Token (TPOT), metrics critical for enterprise-scale AI deployment.
Moreh’s framework is currently the only commercial distributed inference solution tailored for the AMD ecosystem. CEO Gangwon Jo noted that the company intends to lower infrastructure costs by optimizing heterogeneous computing, effectively decoupling high-performance AI services from reliance on specific, monolithic hardware environments. Through its subsidiary, Motif Technologies, the firm is expanding its end-to-end capabilities, signaling a push to capture market share alongside partners like Tenstorrent.




Comments (0)
No comments yet. Be the first!