AI companion apps passed 220 million downloads by July 2025. OpenAI’s reported next step removes the screen and adds cameras and sensors. Less interface, more access.

The smartphone does not have to disappear for its app icon to lose control of the interaction. Hardware can make the assistant less visually intrusive while making it more observationally intensive.

Companion apps, voice models, privacy-oriented clouds, local inference, smart glasses and screen-free devices now converge on one product question: what happens when users no longer have to translate every intention into typed instructions?

Hardware can inherit companionship because software normalized it

Dedicated companion hardware is not trying to create conversational attachment from nothing. A researcher counted 206 AI companion apps in Apple’s App Store and 253 in Google Play by July 2025.

total AI companion-app downloads across Apple’s App Store and Google Play by July 2025

Roughly three-quarters of US teens in one survey said they had used AI as a companion at least once for advice, flirting or deep conversations. These are not merely search queries with friendlier punctuation. They are interactions framed around continuity, personality and disclosure.

Embodiment also predates the current general-purpose-agent push. By 2024, AI-enabled products from Hyodol and Intuition Robotics were already being positioned as sources of companionship and assistance for older adults. Because those products were narrower than today’s proposed general-purpose companions, they show that physical presence can be part of the service rather than decorative packaging.

Those figures establish interest in companion software, not demand for an always-available camera-and-sensor device. Opening an app is reversible; placing a listening and seeing object in a room changes the purchase, consent and trust decision. Yet users who treat an assistant as an ongoing counterpart expose the friction of app discovery, screen unlock and manual context entry. Dedicated hardware can remove that friction, but it cannot inherit the app’s reversible consent model.

Voice, not hardware, makes the screen optional

In November 2025, OpenAI merged ChatGPT’s voice mode into its main chat interface, allowing speech, text and visual responses to coexist instead of placing voice behind a separate mode. ChatGPT could already talk; the change made speaking part of the primary destination.

OpenAI then introduced full-duplex voice models that can listen and speak simultaneously. By handling interruption and overlap, they reduce the interface cost imposed by rigid turn-taking.

Sources describe OpenAI’s first device as a moveable, screen-free smart speaker with a camera and sensors, connected to ChatGPT and intended to act as a humanlike companion. Microphones, cameras, sensors and permitted personal history would acquire the context that a screen no longer mediates.

A screen-free label therefore misses the more important change. ChatGPT’s integrated voice experience can still produce text and visual material, while a physical companion can delegate output to another surface. The assistant becomes the entry point and chooses whether the response should be spoken, shown, stored or acted upon.

OpenAI’s device remains a reported plan rather than a shipped product. The company was still said to be targeting a 2026 unveiling and a 2027 release, leaving demand, reliability and unit economics unproven. Voice integration and simultaneous listening still make persistent access more practical across form factors, even if that particular object fails.

Every ambient interaction becomes a routing decision

An ambient companion must decide which signals to process locally, which requests to send to remote infrastructure and which pieces of personal context to retain. The physical object is only the visible edge of that routing architecture.

More capability generally requires more context and compute. Local processing can reduce exposure and cloud cost, but mobile memory and processing power remain constrained. Cloud processing can run larger models, but it introduces transmission, latency, infrastructure cost and a more complicated chain of custody for personal information.

Apple has described an approximately three-billion-parameter on-device model paired with a larger model on Apple-silicon servers through Private Cloud Compute. Google later introduced Private AI Compute, using protected cloud resources to extend what devices can do.

Apple and Google independently arrived at the same design because device limits and inference bills leave little room for a fixed choice between local and cloud processing. Hybrid systems respond to that constraint without eliminating it: every useful ambient interaction still has to run somewhere.

Barclays projected that inference capital expenditure could surpass training capital expenditure within two years and reach $208.2 billion in 2026. As invocation friction falls and usage rises, product teams must optimize the cost of a completed, trusted interaction rather than model cost alone.

Each contender starts from a different layer of the stack

The participants are not approaching the companion from the same position. Their existing products and infrastructure determine what they can build efficiently and what they are reluctant to surrender.

Player Current position Structural constraint
OpenAI ChatGPT, integrated voice, full-duplex audio and a reported dedicated companion Apple’s lawsuit could complicate hardware hiring and supply relationships before an unproven launch
Apple A local model paired with larger server-side intelligence through Private Cloud Compute Capability must fit device memory, processing and privacy expectations
Google Gemini, an Assistant migration and privacy-oriented cloud inference It must migrate users from a mature assistant without breaking established routines
Meta Smart glasses backed by a rapidly expanding AI infrastructure footprint Wearable regulation, component design and the capital cost of continuous inference
Perplexity Software that routes tasks between local and cloud models It demonstrates routing control without an established companion-hardware position

OpenAI moves outward from the model and conversation. Apple moves inward from the device and its privacy architecture. Google must convert an existing assistant system, while Meta can place intelligence in a wearable interface backed by enormous compute. Perplexity suggests that routing itself can be a product without ownership of the dominant device. Each company’s existing layer shapes the companion it finds economically sensible.

Hardware makes industrial constraints part of interface design

App competition turns on model quality, distribution and software iteration. Physical companions add hiring, sensors, component access, manufacturing relationships, battery rules and supply chains. Those constraints determine the product’s size, mobility, availability and ability to process information locally.

Apple’s lawsuit could complicate OpenAI’s hiring and supply chains. Once model companies become hardware companies, personnel disputes can delay component schedules and release dates.

Policy also becomes product-specific. The European Commission proposed exempting wearable technology from removable-battery requirements, clearing a potential hurdle for Meta’s smart glasses. A battery rule that looks peripheral to an AI model can determine whether an ambient interface reaches the market in its intended form.

Meta also shows how small interfaces rely on industrial-scale infrastructure. Its planned expansion of a Louisiana data-center campus aims for more than five gigawatts of compute and would push total spending beyond $250 billion. Lightweight wearables sit at the edge of a heavyweight system.

Even if models converge, competitors still need component access, local-processing capability, protected cloud infrastructure and manufacturing execution. Those advantages determine whether a polished voice can ship at scale.

Ambient access makes trust a technical requirement

App downloads do not establish demand for an always-on companion with cameras and sensors. OpenAI has demonstrated neither consumer reliability nor viable economics for its reported device. Google pushed its upgrade from Assistant to Gemini on most Android devices past its end-of-2025 target into 2026, showing how difficult even a software migration can become across an established ambient platform.

Local models still confront memory and processing limits, while cloud systems create cost and privacy pressure. Hybrid architectures can allocate work more intelligently, but the allocation itself becomes sensitive: what leaves the device, what remains, what is remembered and which model receives it. A buyer therefore needs to know not just whether an assistant runs “on device,” but which tasks leave it and whether its history can move.

Trust is a functional requirement of the interface, not a brand attribute added after performance. A camera-and-microphone companion becomes more useful as it understands more of the room and the person. The same context that improves assistance raises the cost of an error, unwanted transmission or commercially distorted recommendation.

An assistant embedded in routines, rooms, voice preferences and remembered history can make changing providers harder than downloading another app. Providers can convert personal context into a switching cost even when their models reach rough parity. Portable memory and transparent routing keep the interface contestable; concentration depends on whether users can carry their context elsewhere.

Whoever pays the companion supplies part of its objective

A persistent assistant consumes inference while offering less conventional display inventory than a phone-centric service. Subscriptions, transaction fees, commerce commissions and sponsored responses do more than generate different revenue streams. They give the agent different optimization targets as it mediates decisions.

OpenAI staff have reportedly discussed sponsored content for relevant ChatGPT queries and produced mockups involving sidebars and pop-ups. The company has not launched such an advertising product, but the discussions expose the funding pressure surrounding continuous inference.

A confirmed partnership offers narrower evidence. OpenAI’s deal with Kalshi places FIFA World Cup prediction-market data inside ChatGPT search results. That does not show that responses are secretly advertisements, but it demonstrates how commercially supplied third-party information can enter the assistant experience.

The distinction grows more important as the interface becomes ambient. A conventional search page presents links, placements and labels on a visible surface. A conversational companion can synthesize an answer, choose what to mention and decide whether to act. Commercial influence becomes harder to inspect when the interface produces a recommendation rather than displaying inventory.

Subscriptions face limits on consumer willingness to pay. Commerce can subsidize usage but creates incentives around what the assistant recommends. Sponsorship can lower the apparent price while introducing a second customer whose objectives may not match the user’s. The companion’s incentives depend on who pays, what it discloses and whether the user can leave.

AI’s asset-heavy buildout also carries the lower historical returns associated with capital-intensive businesses rather than the software economics Big Tech investors came to expect. An ambient companion may feel intimate and weightless while carrying the costs of hardware, data centers, energy, networks and depreciation.

Portability decides whether the companion becomes a gatekeeper

An ambient companion observes permitted context, routes each request, selects information and decides when commercial intent is relevant. Replacing it may mean rebuilding memories, permissions, room access and routines rather than simply installing another app.

If users can move that context freely, model and device providers must compete repeatedly for trust. If one provider traps it inside an embodied system, habit and accumulated data can outweigh modest differences in model quality. The concentration threshold is ambient access plus accumulated context without practical portability.

Those 220 million downloads showed that people will try companionship through software. A screen-free companion asks whether they will also invite its cameras, memory and commercial incentives into the room. The object may fade into the background; the gatekeeper does not.