Waleed Atallah of Makora, co-founder and chief executive officer, joins theCUBE
Research hosts Gemma Allen and John Furrier to discuss artificial intelligence,
AI, factories at NYSE Wired. Atallah brings deep expertise in GPU and TPU kernel
optimization, performance engineering and serving open-source models. The
conversation examines GPU supply constraints, the role of kernels versus the
CUDA moat, hardware-agnostic strategies and Makora's approach to delivering fast
cost-efficient inference across diverse accelerators. Atallah emphasizes that
small kernel and performance gains scale to substantial cost savings. They note
that a 2–3% utilization improvement can free the equivalent capacity of
thousands of GPUs in large clusters. They highlight Makora's inference platform,
which delivers faster and lower-cost tokens with open-source models and enables
cost-efficient inference across GPUs and TPUs. They predict consolidation or
strategic partnerships as GPU supply constraints drive providers toward
acquisition or alliances. The discussion addresses data center compute, AI
infrastructure, CUDA performance considerations and tokenomics for model
serving. This episode provides actionable insights for data center operators,
performance engineers and AI infrastructure teams seeking to maximize inference
throughput and cost-efficiency across accelerators.
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open this link to automatically sign into the site.
Register For Cube365 Events Platform
Please fill out the information below. You will recieve an email with a verification link confirming your registration. Click the link to automatically sign into the site.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
You are already logged into TheCUBE Network as
Share
Share clip
The clip will play by default when shared. The entire video and transcript is also accessible via the link.
In this interview from theCUBE + NYSE Wired: AI Factories - Data Centers of the Future, Waleed Atallah, co-founder and chief executive officer of Makora, joins theCUBE's Gemma Allen to discuss how GPU kernel optimization is the hidden lever driving AI economics. With GPU procurement backlogs stretching up to 12 months for the latest chips, Atallah explains why extracting more from existing hardware is now the primary competitive differentiator. He breaks down how GPU kernels — the software layer mapping AI model computations to hardware — determine whether a ...Read more