Huawei's New AI Chip Strategy Reveals How China Is Building an Alternative to Nvidia's Dominance
Huawei's Ascend 960DT neural processing unit, targeting Q1 2027, aims to deliver over 1,100 teraflops and scale to one million NPUs per cluster.
158 articles
Huawei's Ascend 960DT neural processing unit, targeting Q1 2027, aims to deliver over 1,100 teraflops and scale to one million NPUs per cluster.
Jiuwanli raised $100M to build on-device inference chips that run massive AI models locally, eliminating cloud latency and privacy risks on a single chip.
On-device AI inference is replacing cloud processing for security and automation, with models shrunk to under 5% of original size while retaining accuracy.
Neuromorphic chips mimic the brain's efficiency to give robots and wearables real-time AI without cloud dependency, using a fraction of the power of.
OPPO's Find X10 Pro Max will feature a dual-NPU chip offering 51% faster AI responses and 40% better power efficiency for on-device intelligence.
Edge AI chips from BrainChip and Axelera now run on-device inference in production hardware, targeting a market projected to reach $96 billion by 2031.
Three major announcements signal a tipping point for on-device inference: AI is leaving the cloud for edge devices, with 47% of enterprises already.
MediaTek's Dimensity 9600 Pro neural processing unit runs 30-billion-parameter AI models on-device while cutting multi-core power use by 61 percent.
NPUs built into work laptops and IoT devices now run AI tasks locally, boosting speed, privacy, and battery life without relying on cloud servers.
Google's LiteRT-LM lets Android developers run AI models entirely on-device, eliminating cloud costs, network delays, and privacy risks with no per-token.
AMD's Ryzen AI Max+ 395 runs 70-billion-parameter AI models locally using 128GB unified memory, matching desktop CPU performance in a mobile chip.
Apple's S11 chip brings a dedicated neural processing unit to Apple Watch for the first time, enabling on-device AI with 50% more memory bandwidth.
Google's free AI Edge Gallery app runs Gemma 4 entirely on your iPhone offline, with no account or internet required after download.
German companies are moving AI inference off public clouds and into private, hybrid, and edge infrastructure to control sensitive data and meet strict.
Industrial AI is splitting into edge and cloud roles, with on-device inference handling millisecond decisions while cloud systems manage fleet-wide.
Compact AI modules now pack 115 TOPS into a postage-stamp footprint, enabling on-device inference that cuts latency, protects privacy, and works offline.
Qualcomm's next Snapdragon will split AI processing across a CPU, GPU, and neural processing unit, targeting on-device agents and 30B-parameter models.
Big Tech is betting billions on on-device inference, with a $1.35B acquisition and HP, Red Hat, and NVIDIA teaming up to keep AI local.
South Korea is spending $7.2 million on domestic neural processing units from FuriosaAI and Rebellions to power government AI, cutting reliance on Nvidia.
A Japanese researcher now controls WebGPU in llama.cpp, the key tool letting AI run locally on browsers and devices without sending data to the cloud.
Arm's new mobile platform brings AI agents and desktop-class ray tracing to smartphones, rendering just one-eighth of pixels via neural processing units.
Apple's 2027 home sensor will use on-device AI to understand rooms without recording continuous video, pairing with a new home security subscription.
Photonic AI chips that use light instead of electricity are cutting data bottlenecks, with the optical transceiver market projected to hit $112 billion by.
Humanoid robots will need multiple terabytes of local storage to run on-device inference, and that storage bottleneck may matter as much as the chips.
Neural processing units are fueling a mobile AI chip market set to nearly triple to $396 billion by 2036, putting smarter on-device AI in your pocket.
NVIDIA's PAIR tool turns your home network into a local AI cluster, cutting a multi-agent task from 18 minutes to under 9 by sharing work across devices.
MIPS is launching three specialized edge AI platforms instead of one, targeting inference, real-time control, and safety-critical neural processing units.
Taiwan's pharmacies will run AI drug-safety checks entirely on local devices, keeping patient records off the cloud and serving over 4.67 million seniors.
AI chip patents grew 114% in five years, nearly double the broader semiconductor industry, as companies redesign processors from the ground up for.
Your neural processing unit may sit mostly idle on the factory floor if the surrounding software stack isn't ready to use it.
NPUs are moving AI off the cloud and onto your device, with chips hitting 50 TOPS and running 120-billion-parameter models in a shoebox-sized PC.
ChatPPT cut cloud costs by over 50% using on-device inference, processing simple tasks locally while reserving the cloud for complex workloads.
NPUs are now the defining feature of modern chips, powering a SoC market set to hit $413 billion by 2035 as on-device AI becomes standard.
NVIDIA's Jetson Orin Nano 2 will deliver twice the edge AI performance at 40% less power, making on-device inference practical for robots and drones.
Xiaomi's $3.1 billion neural processing unit, the Xring D100, will power autonomous driving in its EVs by 2027, cutting reliance on Nvidia chips.
SOT-MRAM memory completes AI write operations in 2 nanoseconds using just 2 picojoules, making on-device inference fast enough to replace cloud-dependent.
New dual-processor boards combine AI inference with real-time motor control on a single device, cutting latency to milliseconds for robots and industrial.
NXP's "neural axis" splits robot AI across three chip layers, cutting cloud dependency and targeting a robotics market set to hit $16.6 billion by 2030.
Anthropic's universal hardware control system could let Claude AI command factory and lab equipment directly, slashing custom integration work for.
AI chips hit a memory wall as chipmakers embed computing directly into memory systems, a shift backed by billions in new fab investments redefining neural.
Developers are ditching cloud AI bills by running free, local models with tools like Ollama, cutting token costs to zero while keeping data private.
Edge inference can cut daily video uploads from 11 terabytes to mere metadata; here is how to decide which AI workloads belong on-device versus cloud.
Enterprises are shifting AI off the cloud to cut costs, with hybrid on-device inference saving up to $650,000 per 1,000 employees.
Neural processing units are turning laptops into local AI powerhouses, fueling a market set to surge from $4.3 billion in 2026 to $86.2 billion by 2035.
Apple's new M5 Ultra and M6 chips enable local AI inference at 120+ tokens per second, rivaling cloud services while keeping sensitive data on-device.
Wearable Devices is pivoting from gesture controllers to Physical AI, using neural muscle sensors to give robots real-time insight into human intent.
Local AI hardware lets developers ditch cloud API bills forever; Pinea Pi's upcoming edge device runs on-device inference fully offline for a one-time.
Augmented training data boosted robot success rates from 8.3% to 75%, showing why on-device inference and tiered cloud-to-edge AI are key to scaling.
A $60 Arduino board can run a ChatGPT-like AI offline, but its neural processing unit sits completely idle while the CPU handles everything.
A 125M-parameter model now autocompletes piano melodies in real-time on your laptop, with millisecond latency and no cloud connection required.