Apple Partners With Nvidia to Improve Performance Speed of Its AI Models

Apple used Nvidia’s inference acceleration framework for its open-source Recurrent Drafter technique for AI models.

Advertisement
Written by Akash Dutta, Edited by Siddharth Suvarna | Updated: 19 December 2024 17:09 IST
Highlights
  • Apple published a paper on Recurrent Drafter earlier this year
  • Nvidia’s TensorRT-LLM acceleration framework was used for this
  • Apple claims the process resulted in 2.7x faster token generation

Apple had earlier stated the Recurrent Drafter can improve token generation by up to 3.5 tokens per step

Photo Credit: Reuters

Apple is partnering with Nvidia in an effort to improve the performance speed of artificial intelligence (AI) models. On Wednesday, the Cupertino-based tech giant announced that it has been researching inference acceleration on Nvidia's platform to see whether both the efficiency and latency of a large language model (LLM) can be improved simultaneously. The iPhone maker used a technique dubbed Recurrent Drafter (ReDrafter) that was published in a research paper earlier this year. This technique was combined with the Nvidia TensorRT-LLM inference acceleration framework.

Apple Uses Nvidia Platform to Improve AI Performance

In a blog post, Apple researchers detailed the new collaboration with Nvidia for LLM performance and the results achieved from it. The company highlighted that it has been researching the problem of improving inference efficiency while maintaining latency in AI models.

Advertisement

Inference in machine learning refers to the process of making predictions, decisions, or conclusions based on a given set of data or input while using a trained model. Put simply, it is the processing step of an AI model where it decodes the prompts and converts raw data into processed unseen information.

Earlier this year, Apple published and open-sourced the ReDrafter technique bringing a new approach to the speculative decoding of data. Using a Recurrent neural network (RNN) draft model, it combines beam search (a mechanism where AI explores multiple possibilities for a solution) and dynamic tree attention (tree-structure data is processed using an attention mechanism). The researchers stated that it can speed up LLM token generation by up to 3.5 tokens per generation step.

Advertisement

While the company was able to improve performance efficiency to a certain degree by combining two processes, Apple highlighted that there was no significant boost to speed. To solve this, researchers integrated ReDrafter into the Nvidia TensorRT-LLM inference acceleration framework.

As a part of the collaboration, Nvidia added new operators and exposed the existing ones to improve the speculative decoding process. The post claimed that when using the Nvidia platform with ReDrafter, they found a 2.7x speed-up in generated tokens per second for greedy decoding (a decoding strategy used in sequence generation tasks).

Advertisement

Apple highlighted that this technology can be used to reduce the latency of AI processing while also using fewer GPUs and consuming less power.

 

Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.

Further reading: Apple, Nvidia, AI, Artificial Intelligence
Advertisement

Related Stories

Popular Mobile Brands
  1. HMD Pro India Confirmed to Launch in India Soon
  2. Smartphone Screen Protectors Get New BIS Standard in India: What Changes
  3. Vivo V80 is All Set to Launch in India on This Date
  4. Amazon Sale 2026 Will Offer These Smartphones at Discounted Prices
  5. Flipkart Teases Early Bird Offer on iPhone 17 Ahead of Big Billion Sale
  6. Airtel Rs. 999 and Above Postpaid Plans Get iCloud+ Benefits for Free
  7. Amazon Says Festive Sale Prices Are 2026's Lowest for Electronics
  8. Oppo Find X10 Series With Up to 200-Megapixel Cameras Debuts: See Price
  9. Samsung Galaxy S27 Series Could Bring LPDDR6, UFS 5.1 Upgrades
  10. iQOO Z11 Lite 5G Gets New RAM, Storage Variant in India
  1. Samsung Galaxy S27, S27 Plus, S27 Pro, S27 Ultra RAM and Storage Details Leaked Online
  2. Coldcard Whitehats Transfer Stolen 52.37 BTC to Recovery Trust
  3. ChatGPT Gets New Privacy Center With Easy Access to Privacy Controls: Here's How to Use it
  4. WhatsApp Reportedly Brings Third-Party Agent Chats to iOS Beta Users
  5. Oppo Find X10 Pro Max Launched With Triple 200-Megapixel Cameras, Alongside Oppo Find X10, X10 E: Price, Specifications
  6. India to Enforce BIS Registration Rules for Smartphone Screen Protectors From April 2027
  7. HMD Pro India Confirmed to Launch in India Soon; Flipkart Availability Confirmed
  8. X Launches Cashtag Partner Programme With Coinbase, Gemini, Kraken and Other Trading Platforms
  9. Asus, Acer and HP Googlebook 14 Announced Alongside Dell XPS Googlebook, Lenovo Googlebook 15: Price, Features
  10. MediaTek Dimensity CX C10 Max Platform Launched With 55 TOPS NPU for AI-Powered Googlebook Devices
Download Our Apps
Available in Hindi
© Copyright Red Pixels Ventures Limited 2026. All rights reserved.