The first half of 2026 made one thing clear: AI is moving from a chat box into the systems that run work. The rest of the year won’t be defined by one magical model. It’ll be defined by better workflows, tighter controls, cheaper inference, and much less patience for vague AI promises.
I reviewed fresh reporting from Stanford HAI, Gartner, Forrester, McKinsey, the OECD, the European Commission, OpenAI, Anthropic, Google DeepMind, Salesforce, HubSpot, MIT Technology Review, Harvard Business Review, and other research sources. Here’s my H2 2026 outlook for business and tech.
Quick answer: Expect agentic AI to spread through bounded workflows, the EU AI Act to become a day-to-day compliance issue, smaller models to move more work onto devices, and ROI measurement to become a board-level requirement. Healthcare, video, and safety research will make real progress, but mostly in narrow, testable domains.
The 10 AI predictions for H2 2026
The table below is my editorial forecast. Probability means my estimate that the prediction will be visible in business or product news by the end of 2026. It isn’t a survey result.
| Prediction | Probability | Timeline | Business and tech impact |
|---|---|---|---|
| Bounded agentic workflows reach production | 90% | Q3–Q4 2026 | High |
| Enterprise agents need identity, permissions, and observability | 95% | Already underway; accelerates in H2 | High |
| EU AI Act obligations become operational | 95% | August 2, 2026 and after | High |
| On-device AI becomes a default product layer | 85% | H2 2026 | Medium–high |
| Video models move into production previsualization | 90% | H2 2026 | Medium–high |
| AI energy and inference costs constrain scale | 95% | Immediate | High |
| Small open-weight models gain share | 90% | H2 2026 | High |
| Healthcare AI advances through narrow workflows | 85% | H2 2026 | High, but regulated |
| Safety research becomes a deployment loop | 90% | Immediate | High |
| Outcome-based AI ROI standards become normal | 90% | Q3–Q4 2026 | High |
AI use is no longer niche. The OECD’s web-traffic measure rose from roughly 18% of people across GPAI countries in January 2025 to 28% in January 2026, while Microsoft’s separate global telemetry put generative-AI diffusion at 16.3% of the world’s population in the second half of 2025. The measures use different populations and methods, but they point in the same direction: adoption is spreading fast. (OECD; Microsoft AI Economy Institute)
1. Agentic AI workflows will move from demos to bounded production
Agentic AI is software that can pursue a goal, use tools, and take several steps with limited human prompting. In H2 2026, the winning deployments won’t be unconstrained digital employees. They’ll be narrow workflows with clear inputs, permissions, and stop conditions.
The OECD draws a useful line between an AI agent and agentic AI. An agent can perceive and act toward a goal. Agentic AI usually refers to multiple coordinated agents that break down tasks and operate for longer periods with less supervision. That distinction matters because a chatbot that drafts an email isn’t the same thing as a system that checks a claim, retrieves policy rules, updates a case, and escalates an exception. (OECD)
Stanford HAI’s 2026 AI Index shows why this shift is plausible and why it still needs limits. AI agents made a large jump on computer-use benchmarks, yet they still fail a meaningful share of structured attempts. The result is not autonomy everywhere. It’s selective autonomy where a business can tolerate an error and catch it quickly. (Stanford HAI)
The first production workflows will cluster around work that is repetitive and observable:
- Customer-service triage and resolution routing.
- Software issue investigation, test generation, and pull-request review.
- Procurement checks, invoice matching, and internal knowledge retrieval.
- Marketing measurement, audience segmentation, and campaign adjustments.
- Research tasks that gather sources and produce an auditable first draft.
My prediction is simple: by December, “we have an agent” will mean less than “we have one workflow that runs reliably.” Buyers will ask for completion rates, escalation rates, audit trails, and cost per accepted outcome.
2. Enterprise AI agents will become an infrastructure and security problem
Agent infrastructure is the control layer that gives agents identity, access, memory, tool permissions, and observability. It will become a core enterprise architecture category before the end of 2026.
Gartner’s July guidance says agentic systems need orchestration, memory and state management, API governance, and semantic observability. The practical message is blunt: static policies won’t stop a misconfigured agent acting at machine speed. (Gartner)
The risk is easy to picture. An agent connected to email, a CRM, a payments tool, and a knowledge base can do useful work. It can also leak data, retain poisoned instructions, call the wrong API, or repeat a failing loop. The OECD has highlighted prompt injection, memory poisoning, tool misuse, and weak permission design as distinct risks for agentic systems. (OECD)
Forrester’s current “cognitive operating model” work adds an organizational angle. It treats a cognitive skill as the basic unit of work, whether a person, an agent, or a human-agent team performs it. That gives leaders a better question than “Which jobs will AI replace?” Ask which skills should be automated, augmented, or kept human-led. (Forrester)
In practice, enterprises will need at least these controls:
- A registry of every agent, owner, purpose, model, data source, and risk tier.
- Least-privilege access, with explicit approval for high-impact actions.
- Sandboxed tools and circuit breakers that can stop a runaway workflow.
- Tamper-resistant logs for prompts, retrieved context, decisions, and actions.
- Continuous evaluation against real production tasks, not only lab benchmarks.
Salesforce is already framing Agentforce around trusted data, a trust layer, model choice, and Model Context Protocol connections. OpenAI’s July investment guidance makes the same point from another angle: leaders need visibility into usage, spend, products, and models before they scale advanced workflows. (Salesforce; OpenAI)
3. EU AI Act enforcement will become an operating issue in August
The EU AI Act is a risk-based legal framework that assigns different duties to providers and deployers based on how an AI system is used. August 2026 is the point when transparency and broader application duties become a live operating concern for many companies.
The European Commission says the Act becomes fully applicable on August 2, 2026, with exceptions. Prohibited practices and AI-literacy duties already applied from February 2025. General-purpose AI obligations started in August 2025. The Commission’s current timeline also reflects the May 2026 political agreement that extends some high-risk product rules to 2028 and certain high-risk areas to December 2027. (European Commission; EUR-Lex)
The near-term work for businesses is less dramatic than a courtroom headline. It is paperwork, product design, vendor review, and evidence:
- Tell people when they’re interacting with a machine where transparency rules require it.
- Make certain AI-generated text, image, audio, and video outputs identifiable.
- Map general-purpose models and downstream systems across the supply chain.
- Document risk management, logging, human oversight, accuracy, and cybersecurity.
- Train staff who deploy or supervise AI systems.
The European Commission and national authorities are responsible for supervision and enforcement. The Commission’s July 2026 cybersecurity and AI action plan also points toward more evaluation capacity before advanced models reach the EU market.
My prediction is that the AI Act will change procurement before it changes product roadmaps. A vendor that can’t explain its data handling, model updates, incident process, or audit evidence will face longer sales cycles in Europe and, eventually, elsewhere.
4. On-device AI will become a default product layer
On-device AI runs at least part of an AI workload locally on a phone, laptop, car, or other edge device instead of sending every request to a remote data center. In H2 2026, local AI will win on privacy, latency, offline use, and cost.
Apple’s current Apple Intelligence direction combines on-device processing with Private Cloud Compute for harder requests. Apple also says developers can use its Foundation Models framework for offline features without paying per request. Google DeepMind’s Gemma page positions Gemma 4’s smaller E2B and E4B variants for mobile and IoT devices, while its larger variants target personal computers. (Apple; Google DeepMind)
That creates a two-tier pattern. Local models will handle summarization, translation, transcription, image understanding, search over personal files, and simple actions. Cloud models will handle long reasoning, broad research, and high-stakes or compute-heavy work.
This shift will change product design. Teams won’t ask only, “Which model is smartest?” They’ll ask:
- Which tasks must work without a network?
- Which data should never leave the device?
- How much battery, memory, and latency can the feature use?
- When should the device hand off to a larger cloud model?
- Can the user see and control that handoff?
The result won’t be an end to cloud AI. It’ll be a routing problem. The best products will quietly choose the smallest model that meets the quality bar, then escalate only when the task demands it.
5. Multimodal video models will move from clips to production workflows
A multimodal video model generates or understands video while working across text, images, audio, motion, and other inputs. The big H2 shift will be from novelty clips to repeatable previsualization, advertising, training, and interactive media workflows.
OpenAI’s Sora 2 showed the direction: more realistic motion, synchronized dialogue and sound effects, stronger control, and better handling of physical behavior. OpenAI also says the Sora product was no longer available after April 26, 2026, so Sora 2 should be treated as a published capability milestone, not as a currently available product. (OpenAI)
Google DeepMind’s Veo 3 and current Veo 3.1 product materials emphasize native audio, prompt adherence, reference images, character consistency, scene extension, and camera controls. Runway Gen-4 focuses on consistent characters, objects, locations, and world settings across shots. Stanford’s 2026 AI Index reports that video systems are beginning to model how objects behave, not just how they look. (Google DeepMind; Runway; Stanford HAI)
That combination is useful for business because production work is mostly revision. A model that can preserve a character, product, setting, and sound design across multiple shots is more valuable than one that creates a perfect eight-second demo once.
Expect these use cases to grow first:
- Storyboards and previsualization before a live shoot.
- Localized product videos with controlled brand assets.
- Safety training and simulated scenarios.
- Game worlds and interactive marketing experiences.
- Fast concept testing for agencies and in-house creative teams.
The human role won’t disappear. Creative direction, rights clearance, likeness consent, continuity review, and final editorial judgment will matter more as generation gets cheaper.
6. AI energy consumption will become a constraint on growth
AI energy consumption includes the electricity, water, cooling, hardware, and data-center capacity needed to train and run models. In H2 2026, companies will treat it as a cost, infrastructure, and public-policy issue at the same time.
Stanford HAI’s 2026 research-and-development chapter tracks a larger environmental footprint across power, water, and emissions as AI compute expands. The OECD’s compute-and-climate work says policymakers still lack consistent data for measuring national AI compute capacity and its environmental impact. Salesforce is also publishing model-card work that includes environmental metrics. (Stanford HAI; OECD; Salesforce)
The business implication is straightforward: lower token prices don’t guarantee lower total cost. Long-running agents can make many calls, retry, retrieve large contexts, and keep tools active. A small model that finishes a task in one pass may be cheaper and greener than a larger model that looks cheaper per token but needs repeated corrections.
I expect three changes before year-end:
- Finance teams will ask for inference cost by product and workflow.
- Data-center plans will face more scrutiny from utilities, communities, and regulators.
- Model routing, caching, quantization, and smaller open-weight models will become normal optimization tools.
Energy won’t stop AI adoption. It will decide which deployments deserve to scale.
7. Small open-weight models will take more work away from the frontier cloud
An open-weight model is a model whose trained parameters are released so developers can run, adapt, or fine-tune it under the model’s license. That isn’t always the same as fully open-source AI, because training data, code, and every component may not be public.
Small models will gain ground because most business tasks don’t need the largest model available. They need a predictable model that can classify a ticket, extract fields, summarize a call, translate a message, or trigger a safe tool call at low cost.
Stanford HAI’s 2026 Index says open development continues to scale and that smaller, carefully trained systems can perform strongly against much larger models on selected tasks. Google DeepMind’s Gemma 4 lineup gives developers a concrete set of open models for phones, laptops, and cloud deployments. Microsoft’s AI Diffusion report also points to DeepSeek’s open release and low-cost access as major drivers of adoption in markets that were underserved by traditional providers. (Stanford HAI; Google DeepMind; Microsoft AI Economy Institute)
The practical advantage is control. A company can run a smaller model inside its network, tune it on domain language, keep data local, and switch vendors more easily. The trade-off is that the company owns more of the evaluation, monitoring, security, and update burden.
By December, “best model” comparisons will look less like a single leaderboard. They will look more like a portfolio:
- A tiny local model for fast, private tasks.
- A mid-size open model for domain workflows.
- A frontier API for ambiguous or high-stakes reasoning.
- A deterministic rules engine for actions that must never drift.
8. Healthcare AI will advance through narrow, validated workflows
Clinical AI is AI used in care delivery, medical operations, research, or patient communication. The strongest healthcare advances in H2 2026 will come from narrow workflows with human review, not from replacing clinicians with a general-purpose chatbot.
Stanford HAI’s 2026 medicine chapter highlights virtual cell models, AI-generated clinical notes, medical imaging, diagnostic orchestration, and open medical models. Google DeepMind’s Gemma ecosystem includes MedGemma for medical text and image comprehension. The OECD’s review of EU AI deployment also identifies diagnostics and hospital operations as practical health use cases, while warning that fragmented data and trust still limit scale. (Stanford HAI; Google DeepMind; OECD)
The first winners will reduce friction around care:
- Ambient documentation that turns visits into draft notes.
- Imaging triage that moves urgent cases up a queue.
- Scheduling and bed-occupancy forecasting.
- Patient instructions translated into clear, local-language guidance.
- Research systems that connect papers, protocols, and candidate hypotheses.
The bar is higher than “the model sounds convincing.” Clinical systems need validation on representative populations, clear escalation rules, traceable evidence, privacy controls, and post-deployment monitoring. A system can be useful without being autonomous.
My forecast is a lot of modest wins that add up: less documentation burden, faster triage, better research search, and more targeted clinical decision support. The headline won’t be “AI replaces the doctor.” It’ll be “AI removes one more bottleneck from a care team.”
9. AI safety and alignment research will become part of the release cycle
AI alignment means making an AI system behave in ways that match human goals, rules, and safety constraints. In H2 2026, alignment work will look less like a separate ethics program and more like continuous testing, monitoring, and incident response.
Stanford HAI’s responsible-AI chapter says capability reporting is far ahead of responsible-AI reporting. It also finds that safety, fairness, privacy, and accuracy can trade off against each other. That makes one benchmark score a poor proxy for safe deployment. (Stanford HAI)
Anthropic is pushing mechanistic interpretability and model-behavior research, while its July jailbreak framework proposes a shared severity scale for cyber jailbreaks. OpenAI’s July bio bug bounty extends external testing for universal jailbreaks in frontier models. The OECD is calling for shared testing environments and information-sharing around prompt injection, agent vulnerabilities, jailbreaks, and model poisoning. (MIT Technology Review; Anthropic; OpenAI; OECD)
The next step is operational. Model releases will increasingly ship with:
- Capability and misuse evaluations.
- Adversarial red-team results.
- Access tiers for sensitive capabilities.
- Monitoring and rapid-remediation plans.
- Clearer third-party testing and disclosure channels.
This won’t make advanced models perfectly safe. It will make safety claims easier to compare and failures easier to investigate. That is real progress.
10. Outcome-based AI ROI will replace vanity metrics
Outcome-based AI ROI measures the cost and value of an accepted business result, not just prompts, tokens, logins, or hours of experimentation. This will be the most practical AI trend for CFOs and operating leaders in the second half of 2026.
OpenAI’s July guidance tells leaders to measure useful work per dollar and cost per accepted outcome. Forrester’s 2026 marketing measurement work says AI can shorten the gap between an insight and an action, but still needs human oversight. Gartner’s finance guidance says AI shifts accountability upstream into data, assumptions, model choice, and system design. HBR is publishing similar work on new performance metrics for employees, AI systems, and combined output. (OpenAI; Forrester; Gartner; HBR)
A serious ROI dashboard should track at least five layers:
- Adoption: Who uses the system, for what work, and how often?
- Quality: How often is the output accepted without correction?
- Speed: How much cycle time or queue time does the workflow remove?
- Economics: What does each accepted outcome cost, including retries, tools, review, and infrastructure?
- Risk: What errors, privacy events, escalations, or revenue losses does the system create or prevent?
This changes the funding conversation. A pilot that saves time but creates review work may not be a win. A workflow with a higher model bill may be a better investment if it resolves more cases, reduces rework, and produces evidence an auditor can trust.
By year-end, leading teams will fund workflows that compound. They will build shared identity, data, evaluation, and observability layers once, then reuse them across many use cases. The unit of investment will be a measurable workflow, not a vague “AI transformation” program.
What leaders should do before 2026 ends
The safest H2 strategy is to scale a few valuable workflows while building the control plane around them. I’d prioritize these five moves:
- Pick one workflow with a clear owner, baseline, and failure cost.
- Give the agent the least authority it needs, not the most authority it can use.
- Test the workflow on real edge cases before expanding access.
- Track cost per accepted outcome and review it with finance every month.
- Map EU AI Act duties, vendor evidence, data flows, and human-oversight requirements now.
The big story of the rest of 2026 won’t be that AI became human. It will be that businesses got better at deciding where AI should act, where it should assist, and where it should stop.
Sources
- Stanford HAI, 2026 AI Index Report
- Stanford HAI, Technical Performance
- Stanford HAI, Economy
- Stanford HAI, Research and Development
- Stanford HAI, Responsible AI
- Stanford HAI, Medicine
- OECD, Agentic AI concepts and adoption
- OECD, Collective AI security
- OECD, GenAI chatbot usage
- OECD, AI compute and the environment
- Gartner, infrastructure for agentic AI
- Gartner, emerging GenAI adoption trends
- Gartner, AI-enabled forecasting
- Forrester, AI cost management
- Forrester, cognitive operating model
- Forrester, agentic development platforms
- Forrester, AI and marketing measurement
- European Commission, AI Act
- European Commission, AI Act governance and enforcement
- EUR-Lex, Regulation (EU) 2024/1689
- OpenAI, managing AI investments in the agentic era
- OpenAI, GPT-5.6
- OpenAI, Sora 2
- OpenAI, Bio Bug Bounty
- Anthropic, Claude Sonnet 5
- Anthropic, jailbreak severity framework
- Google DeepMind, Veo
- Google DeepMind, Gemma
- Runway, Gen-4
- Apple, Apple Intelligence and Siri
- Microsoft AI Economy Institute, Global AI Adoption in 2025
- Salesforce, AI at Salesforce
- HubSpot, 2026 State of Marketing Report
- Harvard Business Review, Technology and analytics
- MIT Technology Review, AI topic
- IDC AI Economy forecast hosted by Salesforce
- World Economic Forum, Future of Jobs 2025


