top of page

Search Results

173 results found with an empty search

  • July 15, 2026: The Professionals Protecting Their Focus Time Stopped Doing It Manually

    The productivity problem most senior professionals have is not that they lack good AI tools. It's that their calendar fills up faster than they can manually protect time to use them. You've probably tried the Sunday-night ritual: block two-hour focus windows across the week. By Tuesday at 11am, one is gone to a scheduling conflict, another got voluntarily sacrificed for a "quick call," and the third never felt legitimate enough to defend in the first place. The problem isn't discipline. It's that protecting attention requires a daily decision you're making against a moving target, and you're making it manually every single day. Reclaim.ai approaches this differently. It connects to your Google Calendar and automatically inserts and defends focus blocks, task time, and personal habits around your existing meetings, without you touching a thing each morning. In this post. Why Manual Time-Blocking Keeps Failing, the structural reason your Sunday blocks don't survive the week What Reclaim Actually Does, how the scheduling layer works without requiring new tools or habits What a Meeting-Heavy Tuesday Looks Like When This Is Running, the concrete, felt difference in your day Who Gets the Most From It, and Who Doesn't, an honest fit assessment for different calendar situations Try It This Week, immediate steps to test it this week Manual Time-Blocking Keeps Failing Because It's a Static Answer to a Dynamic Problem When you block 2–4pm as "focus time," you're creating a calendar entry with no intelligence. It doesn't know that the meeting at 1:30pm ran long. It doesn't know a deliverable just became urgent. It doesn't re-sequence itself when a morning meeting gets added and compresses your pre-lunch window to 20 minutes, too short for real work, too long to skip. The deeper issue is decision fatigue. Each morning, effective time management requires you to assess your task list, your meeting load, your energy, and your deadlines, then manually move blocks around to find a viable focus window. That decision, repeated daily across a dynamic calendar, consumes cognitive resources before you've done a single minute of actual work. Practitioners who've moved off manual blocking describe a consistent pattern: the schedule looked reasonable in advance, felt wrong by mid-morning, and required mental re-planning that consumed the exact bandwidth they were trying to protect. Reclaim Treats Your Calendar as a Living Schedule, Not a Static Grid Reclaim.ai connects to Google Calendar, currently the primary supported platform, and acts as a scheduling layer that runs continuously in the background. You set your preferences once. It works from there. The core mechanics, in plain terms: Focus time blocks. Reclaim identifies open windows in your calendar and automatically inserts focus sessions based on how many hours you've told it you need per week. When a meeting gets added and would eat into a focus block, Reclaim attempts to move the block to another open window rather than simply losing it. Task scheduling. You add tasks with deadlines and rough time estimates, and Reclaim schedules them into available slots, weighted by deadline and priority. Think of it as a task list that also holds its own calendar space rather than sitting in a separate app you have to manually translate into time. Habit protection. You designate recurring commitments, lunch away from your desk, a walking break, an end-of-day review, and Reclaim treats these as soft but defended blocks. They yield to urgent meetings but otherwise hold their position. The result is a calendar that reflects your actual working priorities, not just your meeting obligations. Practitioner coverage highlights the tool's core value as reducing the daily "when can I actually do this?" decision to near-zero. The calendar answers that question automatically. Action step. Before testing anything, spend five minutes listing the three types of time you most need protected each week, deep thinking blocks, task completion windows, or recovery habits. That's the input Reclaim needs from you. Everything else it handles. What a Meeting-Heavy Tuesday Looks Like When This Is Running The felt experience matters more than the feature list. You have four meetings on a Tuesday: 9am, 11am, 2pm, and 4pm. Left to a standard calendar, your day looks like a series of 45-to-90-minute fragments between those anchors, none quite long enough for focused analytical work, all requiring a mental context switch on entry and exit. With Reclaim active: The 90-minute gap between the 9am and 11am meetings gets flagged as viable focus time if you've set a preference for morning deep work. A defended block appears automatically. A task you've entered, say, a document review due Thursday, gets scheduled into the post-2pm window rather than floating as an unscheduled obligation. Your 12:30pm lunch habit holds unless someone books a conflict, in which case Reclaim shows you the conflict rather than silently dropping your habit. You arrive at your desk knowing where your focus time lives that day. You didn't build that schedule. The system did. Practitioner coverage from 2026 highlights this as the practical differentiator: not any single feature, but the aggregate reduction in daily schedule-management overhead for professionals carrying eight or more calendar events per day. Who Gets the Most From It, and Who Doesn't Reclaim is a strong fit for a specific calendar situation. It is not for everyone. Gets the most value. Senior ICs or managers with 6–10 or more scheduled meetings per week, where deep work time exists but is fragile and easily overwritten Professionals living in Google Calendar as their primary scheduling surface Anyone who has tried manual time-blocking and finds it doesn't survive real-week conditions Weaker fit. Professionals with primarily self-directed calendars and few external meeting demands, the tool solves a fragmentation problem that doesn't exist for them Teams on Microsoft Outlook or other calendar platforms: Reclaim's primary integration is Google Calendar, which limits its usefulness in Outlook-heavy organizations Anyone looking for a broader AI agent or full task management overhaul, Reclaim is a narrow, calendar-specific layer, not a complete productivity system One practical note: Reclaim is a cloud tool. Your calendar data, task details, and schedule preferences are processed on its servers. For most professionals, this is a reasonable exchange, calendar data is already in Google's infrastructure. If your role involves particularly sensitive scheduling information, review Reclaim's data handling terms before connecting. A realistic setup for many professionals combines tools by task type: cloud tools like Reclaim for calendar intelligence, and enterprise-grade platforms (Google Workspace Gemini, for those whose employers provide it) for document and communication work involving confidential content. The two serve different problems and don't need to conflict. Try It This Week Audit your last two weeks of calendar data before setting up anything. Count how many focus blocks you manually created versus how many survived intact to their intended purpose. If the survival rate is below 50%, you have a fragmentation problem that automation is likely to help. Start with one preference, not three. When you first configure Reclaim, set a single goal, five hours of focus time per week, for example, and let it run for two weeks before adding task scheduling or habit blocks. Layering too many preferences at once makes it harder to see what's actually working. Log what you do in the first few defended blocks. The goal isn't just to have protected time on your calendar. It's to confirm you're using it for work that requires sustained thinking. If you fill defended blocks with email, the system is working but the habit isn't. Fix the habit, not the tool. If you're on Google Workspace through your employer, check whether your IT function has any restrictions on third-party calendar integrations before connecting Reclaim. Most organizations allow it; some don't. A two-minute check saves a later conversation. When did you last finish a workday feeling like you had adequate time to actually think? If the answer takes longer than three seconds, that's the signal, and the calendar is where the fix starts. The most durable productivity edge isn't a smarter chat tool. It's protecting the time to use your judgment before someone else fills it with a meeting request. If you want to stay current on what AI means for individual professionals, the practical tools and approaches that actually change how your days work, not the organizational hype, Personal Agenticism is where those insights live. Subscribe at Agenticism on Substack for the curated weekly delivery. Sources Reclaim.ai, View Article Top 5 AI Scheduling & Calendar Tools 2026, Deepak Gupta, View Article Reclaim AI Review, Lifestack, View Article Reclaim AI Automation Tools Blog, View Article

  • The Hidden Tax on Your AI Agents: How Retry Loops Are Quietly Draining Up To 60% of Your AI Budget

    One company's overnight AI bill hit $72,000, not from a massive new deployment, but from a single agent stuck in a retry loop, repeatedly calling the same tool until someone noticed the charge. That incident, documented by budget-limit tooling provider SatGate in April 2026, is not an outlier. It is what happens when billing systems designed for simple, single-turn AI queries meet the messier reality of autonomous AI agents that can fail, retry, and fail again, on your dime. If your organization is running AI agents in any live business process, you are almost certainly paying a hidden tax you cannot yet see on your invoices. A Gartner survey of 180 AI-mature enterprises, conducted in Q4 2025, found that 67% cite unpredictable agent token costs as their top barrier to moving AI from pilot to full production. The billing model most AI providers use was not built for how agents actually behave, and the gap between what you expect to pay and what you actually pay is widening fast. The Trend in Plain Sight To understand the problem, a quick explanation of how AI billing works is useful. Most AI providers charge by the "token," which is roughly a word or part of a word. Every time your AI processes text or generates a response, you pay for the tokens used. Think of it like paying for a phone call by the word spoken, not by the minute. AI agents, though, are different from a simple question-and-answer interaction. An agent is an AI system that takes a goal, breaks it into steps, uses tools (like searching a database or sending an email), checks its own work, and tries again if something goes wrong. Each of those steps, retries, and self-checks burns tokens. When an agent hits an error and retries five or ten times by default, you pay for every attempt. LangChain, one of the most widely used tools for building AI agents, added retry tracking in February 2025 and found that early users were spending 40 to 60 percent of their total AI budget on failed or retried steps. New Relic, an independent software monitoring company, observed a 2.8x average token multiplier in live production agent deployments across 120 customers, compared to what those same tasks would cost in a simple single-turn interaction. Datadog's 2026 State of AI Engineering report found that token usage per request more than doubled year-over-year for the median customer, and quadrupled for the top 10 percent of users. The problem has a name in engineering circles. Industry analysts are now calling it "token debt", the accumulated cost of tokens consumed by failed attempts, verification loops, and parallel sub-agents checking each other's work. In multi-agent systems where several AI instances collaborate on a task, there is also a "swarm tax", which is the overhead of agents coordinating, passing context back and forth, and re-running steps when one agent's output does not satisfy another's verification check. Passing failure context on retries, rather than starting fresh each time, can reduce this consumption by 40 to 60 percent, according to analyses published in mid-2026. Who is moving first. Financial services firms are leading because unpredictable spend on regulated workloads is a compliance and audit problem, not just a budget problem. Healthcare organizations are pushing toward self-hosted agents because of strict rules around patient health data leaving their own systems. Professional services firms are adopting cost-tracking tools fastest because their margins are thin and AI spend is directly visible on client project budgets. Why This Is Happening Now Three things changed between 2023 and 2025 that created this specific problem. First, agents became the default deployment pattern. Two years ago, most enterprise AI was question-and-answer. A user asks, the AI responds, done. Today, organizations are deploying agents that run autonomously over minutes or hours, making dozens of tool calls and self-corrections. The billing infrastructure was never updated to match. Second, default settings in agent frameworks were built for reliability, not cost. Most agent-building tools ship with retry settings of five to ten attempts per failure. That made sense when AI was used in low-volume experiments. At production scale, with hundreds of agents running simultaneously, those defaults become a cost multiplier that compounds silently. Third, providers had no financial incentive to surface the problem. Model providers charge per token. More retries mean more tokens mean more revenue. There was no structural pressure on OpenAI or Anthropic to build billing transparency that would help customers spend less. That pressure is now arriving from enterprise procurement teams who are seeing invoices that do not match their forecasts. Consider a useful analogy. Imagine hiring a contractor who charges by the hour, where the contract says nothing about what happens if they make a mistake and have to redo work. You assumed they would get it right the first time. They assumed retrying was just part of the job. Nobody wrote down who pays for the do-overs. That is the current state of enterprise AI agent billing. Microsoft Azure customers pushed hard enough that Azure added policy controls in preview in July 2025, allowing per-session token limits. OpenAI introduced explicit agent message caps for Enterprise and Business tiers in April 2026, along with separate usage-based billing for its Codex coding agent. These are early responses to a billing friction that has now reached the contract negotiation stage. Key Numbers at a Glance 40–60% of total AI spend from failed or retried steps, reported by early adopters of LangChain's retry telemetry module (LangChain, February 2025) 2.8x average token multiplier in live production agent deployments versus single-turn baselines, across 120 monitored customers (New Relic, August 2025) 67% of AI-mature enterprises cite unpredictable agent token costs as their top barrier to full production rollout (Gartner, Q4 2025) $72,000 overnight billed to one organization from a single runaway retry loop in a production agent (SatGate, April 2026) 41% reduction in monthly OpenAI spend at one fintech company after Datadog surfaced and disabled retry loops in customer-support agents (Datadog, 2025) 35% reduction in effective spend by pilot users who added early termination logic using Weights & Biases prompt tracing (Weights & Biases, May 2025) 20–40% savings reported by organizations using AI FinOps platforms with automated waste detection in multicloud environments (Finout, June 2026) Here's Where This Points Current patterns and the documented scale of retry waste make three outcomes increasingly likely over the next two to three years. By the end of 2026, retry cost attribution will become a standard procurement requirement. The Gartner finding, the $72,000 overnight incident, and Azure's policy controls all point toward enterprise buyers demanding contractual language around retry caps before signing new AI contracts. Providers who cannot offer this will lose deals to those who can. OpenAI's April 2026 agent caps are an early signal that the market is already forcing this change. By 2027, the billing opacity problem will accelerate migration away from per-token APIs for high-volume agent workloads. If the documented 3 to 5x cost differentials in production agent workloads continue, and if open-weight models (AI models whose inner workings are publicly shared, allowing companies to run them on their own systems without per-use fees) continue improving in quality, enterprises running large agent fleets will increasingly move those workloads to self-hosted infrastructure where they control the retry logic entirely. Financial services and healthcare will lead, for cost and compliance reasons simultaneously. AI FinOps, the discipline of tracking, attributing, and optimizing AI spending across an organization, will become a dedicated function inside most large enterprises by 2028. The tooling is already emerging. Datadog, Weights & Biases, Finout, and newer platforms like Cloudgov.ai are building automated detection and remediation. The question is not whether this function will exist, but whether organizations build it proactively or reactively after a billing incident forces the conversation. What This Means for the Budget Owners If you own the budget that covers AI spend, you are currently flying partially blind. The line item on your invoice that says "AI API usage" does not tell you how much of that spend was productive work versus failed retries. An insurance company using Weights & Biases tracing attributed $47,000 in monthly retry spend to a single tool-calling agent, and had no idea until they looked. That kind of invisible waste is almost certainly present in any organization running agents at scale. Your team's standard cost controls do not map cleanly onto this problem. You cannot cap AI spend the way you cap software licenses, because the billing is usage-based and agents can generate usage autonomously. A budget set for 10,000 tokens per day can be consumed in minutes by one stuck agent. AI spend needs the same monitoring infrastructure you apply to cloud computing costs. Cloud bills became unmanageable in the 2015-2020 period until FinOps tools and disciplines emerged to bring them under control. AI agent spend is at the same inflection point now, and the organizations that build the monitoring function early will have significantly better cost predictability when agent deployments scale. For smaller teams without dedicated FinOps resources, the near-term action is simpler. Audit your current agent configurations for default retry settings, and set hard token caps per agent run. A logistics firm that did this using LangChain's retry telemetry cut waste from 58% to 12% of total spend within six weeks. Practical Next Steps In the next 30 days. If you are running agents in production, pull your last 90 days of AI billing data and look for usage spikes that do not correspond to business activity spikes. Overnight charges, weekend surges, or single-day anomalies are the fingerprint of runaway retry loops. If your current provider's billing export does not show this level of detail, that is itself important information for your next contract conversation. In the next 60–90 days. Evaluate one observability tool that provides span-level token attribution, meaning it shows you the cost of each individual step inside an agent run, not just the total. Datadog's AI Observability features, Weights & Biases prompt tracing, and Finout's agentic cost allocation are all documented options. Even a pilot covering your highest-volume agent will surface data that changes how you think about the spend. For larger teams, assign someone to own AI spend attribution as an explicit responsibility. This does not require a new hire. It requires naming the function and giving it teeth in the budget process. For smaller teams, set hard token caps per agent session in your agent framework configuration. Most frameworks support this; most organizations have not turned it on. LangGraph added default checkpointing in April 2025 specifically to prevent infinite retry loops. Vercel's AI SDK introduced configurable retry budgets in January 2025. These controls exist. Use them. When renegotiating AI contracts. Ask providers directly for retry attribution in billing exports and contractual caps on retry-induced overages. Azure has this in preview. The fact that you are asking signals to providers that the market is moving, which is the only pressure that changes billing model design. The Second-Order Story The billing opacity problem is not just a cost management headache. It is quietly reshaping which AI vendors enterprises will trust with their largest workloads, and that shift has consequences that run well beyond the organizations paying the bills. When cloud computing bills became unpredictable in the mid-2010s, enterprises did not just buy better monitoring tools. They renegotiated contracts, moved workloads to reserved capacity, and built internal cloud teams that reduced their dependence on the most expensive managed services. The same dynamic is now starting with AI. Retry cost unpredictability is not just a billing problem, it is a trust problem, and trust problems change buying behavior. For OpenAI and Anthropic, the exposure runs deeper than it first appears. Both companies generate a substantial share of their revenue from usage-based API billing. Enterprise customers running large agent fleets are exactly the high-volume, predictable accounts that anchor that revenue model. When those customers discover that 40 to 60 percent of their spend is retry waste, their first response is to add controls. Their second response, if the controls are insufficient or unavailable, is to evaluate whether they can move to a vendor that provides financial operations data along with better controls, or run comparable models on their own infrastructure and eliminate the per-token exposure entirely. Open-weight models that companies can run on their own systems have been improving steadily in quality, and the cost differential for high-volume workloads is already documented at 3 to 5x. Retry unpredictability accelerates the math that was already pointing toward migration. The companies positioned to capture the displaced spend are the ones offering controllable inference at lower cost. CoreWeave, Together AI, and Fireworks.ai are building infrastructure specifically for enterprises that want to run AI on their own terms. Databricks Mosaic AI and Hugging Face are winning on the deployment of open-weight models with built-in controls. The AI FinOps platforms, Datadog, Finout, Cloudgov.ai, and others, are building a new software category that did not exist two years ago and will likely be a standard enterprise tool by 2028. There is also a talent implication that most organizations have not yet registered. The engineering skills most in demand are shifting from "how do I build an agent" toward "how do I make an agent cost-efficient at scale." Inference optimization and AI cost engineering are becoming distinct specializations, and the organizations that hire for them now will have a structural advantage when agent deployments scale across departments. What Could Slow This Down The billing transparency problem will not resolve quickly, for several reasons. Model providers have a direct revenue interest in maintaining per-token billing without retry attribution. More tokens billed means more revenue. The pressure to change is coming from enterprise procurement teams, not from inside the providers, and procurement pressure moves slowly through multi-year contracts. Enterprise procurement teams currently lack standardized contract language for retry cost caps. Every negotiation is starting from scratch, which slows adoption of the new controls that do exist. Until industry bodies or large buyers establish template language, this will remain a friction point. Observability tooling is still maturing. Datadog, Weights & Biases, and New Relic have built meaningful attribution capabilities, but full coverage across all major AI APIs remains incomplete as of today. Organizations running agents across multiple providers face a more complex attribution problem than those using a single API. Internal skills gaps are also a real constraint. Setting token caps, configuring early termination logic, and interpreting span-level cost data requires engineering knowledge that most operations, finance, and HR teams do not have in-house. The tooling is becoming more accessible, but there is still a gap between "this control exists" and "your team can implement it without dedicated engineering support." Finally, regulatory attention has not yet reached AI billing mechanics. Data privacy and AI fairness are getting regulatory scrutiny. Billing transparency is not, which means there is no external forcing function pushing providers to change faster than their customers can push them. Bottom Line By 2027, AI agent billing will look meaningfully different from today, driven by enterprise pressure that is already visible in contract negotiations and provider product updates. Organizations that build retry attribution and token cap controls into their agent deployments now will spend 30 to 40 percent less on the same workloads than those that wait for providers to solve the problem on their behalf. The providers most exposed are those whose revenue depends on high-volume enterprise API usage without offering the cost controls those enterprises are now demanding. The organizations with the most to gain are those that treat AI spend as a managed cost category today. Challenge your vendors now to help expedite the delivery of visibility, and spend controls. Start exploring alternatives now. Sources OpenAI, Enterprise billing export updates adding separate retry and tool-call categories after customer complaints about agent loops (March 2025). Signals that retry-driven costs were significant enough to require separate reporting. Anthropic, Agent pattern guidance for Claude 3.5/4 recommending explicit stop conditions after observed 3–7x token inflation in multi-step workflows (June 2025). First major provider to acknowledge retry overhead in public documentation. LangChain, Retry telemetry module release; early adopters reported 40–60% of total spend from failed or retried steps (February 2025). Quantifies the hidden tax inside the most widely used agent-building framework. Datadog, AI FinOps dashboard launch tracking "agentic token waste" across OpenAI and Anthropic calls (September 2025); expanded Agent Observability with span-level cost breakdowns and 2026 State of AI Engineering report finding token usage per request more than doubled year-over-year for median customers (July 2026). https://www.datadoghq.com/blog/making-agentic-token-costs-visible-in-production/ New Relic, Observed 2.8x average token multiplier in production agent deployments versus single-turn baselines across 120 customers (August 2025). Independent telemetry confirming the scale of retry overhead. Gartner, Survey of 180 AI-mature enterprises finding 67% cite unpredictable agent token costs as top barrier to production rollout (Q4 2025). Establishes this as a market-wide constraint, not an edge case. Weights & Biases, Prompt and agent cost tracing enabling an insurance company to attribute $47,000 in monthly retry spend to one specific agent; pilot users cut effective spend 35% by adding early termination (May 2025). LangGraph, Default checkpointing introduced to reduce infinite retry loops after reported cases exceeding 10,000 tokens per failed agent trajectory (April 2025). Microsoft Azure OpenAI, Policy controls in preview allowing per-session token limits after enterprise customers requested contractual caps on retry-induced overages (July 2025); sustained in 2026 documentation with automatic quota scaling. Vercel AI SDK, Configurable retry budgets introduced after user reports of runaway costs in production agents (January 2025). OpenAI, Enterprise and Business tiers introduced Agent Mode message caps and separate usage-based billing for Codex agentic workflows (April 2026). https://www.gosearch.ai/faqs/chatgpt-enterprise-pricing-explained-cost-tiers-hidden-fees-gosearch-comparison/ Finout, AI FinOps platform with agentic-specific features including automatic waste detection and reported 20–40% savings in multicloud environments; analysis of how FinOps must evolve for the agentic era (June 2026). https://www.finout.io/blog/how-finops-must-evolve-for-the-agentic-era-of-ai TrueFoundry, AI cost optimization strategies including circuit-breaker controls targeting retry loops in production (June 2026). https://www.truefoundry.com/blog/ai-cost-optimization-strategies Zuplo, Circuit-breaker and per-agent token budget controls for retry loops in production deployments (April 2026). https://zuplo.com/blog/rate-limit-ai-agents-beyond-request-counts SatGate, Documented real-world runaway retry loop incident exceeding $70,000 overnight, driving demand for hard budget caps (April 2026). https://satgate.io/blog/how-to-add-budget-limits-to-openai-api-calls dsanchezcr analysis, Token debt and swarm tax phenomena in agent systems; passing failure context on retries reducing consumption 40–60% (2026). https://dsanchezcr.com/blog/token-debt-finops-agentic-engineering Cloudgov.ai, Agentic AI FinOps platform with automated remediation of retry waste in real time (2026). https://cloudgov.ai/ Technical readers can find detailed customer metrics and benchmarks in the original announcements linked above.

  • July 14, 2026: BCG Found 42% of AI Users Save a Full Day Per Week. Two-Thirds Get No Direction on What to Do With It.

    In this post. BCG's survey of over 11,000 global employees on AI time savings and the guidance gap New research on how AI is reshaping career trajectories for workers 55 and older What code quality data tells us about AI-assisted engineering in 2026 Three research studies arrived this week with a shared thread running through each. AI is delivering real time savings and real disruptions at the worker level. Organizations are largely watching. A BCG survey of over 11,000 global employees found that 42% of employees who use AI regularly save a full workday or more per week. That is a structural shift in available capacity, not a rounding error. The following number explains why most organizations aren't capturing it. Per the same BCG consulting-firm survey, 66% of those employees receive little or no guidance on what to do with the time they save. Nearly seven in ten workers are absorbing freed-up hours they have no instruction for. Time that isn't redirected intentionally fills itself with meetings, ambient distraction, and whatever tasks were next in queue. The tools are working. The operating model around them hasn't been updated to match. Older Workers in AI-Exposed Roles Are Leaving Faster Than They Used To Research from the Center for Retirement Research at Boston College sharpens the picture in a different direction. Workers aged 55 and older in AI-exposed industries are leaving their jobs more often than they did before AI tools became widespread. Before ChatGPT entered professional workplaces, older workers in those same roles were significantly less likely to leave than their younger peers. That pattern has reversed. The Boston College researchers point to automation as a likely driver. Roles may be narrowing or disappearing, or the job insecurity is sufficient that older workers exit before being pushed out. For workers who planned to continue into their late fifties and sixties, this is not only a career disruption. It is an economic one, and the research does not distinguish between voluntary exit and forced transition, which means the true picture may be harder than the numbers suggest. For workforce planners and HR professionals, the operational question is whether your organization's attrition tracking is disaggregated enough to surface this pattern. If AI-exposed roles are losing experienced workers faster than they were two years ago, you are losing institutional knowledge alongside the efficiency gains you are counting. The two outcomes are not independent of each other. AI-Generated Code Carries More Defects, and Developer Trust Has Dropped A CodeRabbit analysis of 470 open-source pull requests found that AI-generated code carries 1.7 times more defects than code written by humans. Stack Overflow's 2025 Developer Survey found developer trust in AI tool accuracy fell to 33%, down from 43% the year before. GitClear's maintainability research found copy-pasted code has nearly doubled since 2022. Three data sets, one dynamic. Engineers are shipping code faster with AI assistance. That code is arriving with more bugs. Developers are growing more skeptical of what the tools produce. And the rising volume of duplicated code is building a maintenance burden that compounds over time. None of this argues for removing AI from engineering workflows. It argues for code review standards that are proportionate to what AI tools actually produce. If your review processes were calibrated before AI-generated pull requests became routine, they were calibrated for a different defect rate, and that gap is now measurable. CodeRabbit provides code review software, which gives it a commercial interest in findings that favor more rigorous review processes. The pull request analysis is publicly available, but that context is useful when weighing the conclusions. The Structure That's Missing Each of these three studies surfaces a version of the same gap. AI tools are producing real outputs at the point of use: time freed, roles shifted, code written. What isn't keeping pace is the organizational layer that determines what happens next. BCG's finding on guidance is the most directly addressable. If your organization has deployed AI tools with any meaningful reach, a significant share of your workforce is likely already saving time with no clear direction on how to reinvest it. Building that second half of the equation is a management decision, not a technology one. Act on These Now Map where saved time is actually going. If AI tools are producing time savings in your function, find out what those hours are being used for. Brief surveys, one-on-ones, and workflow observation can surface whether teams are reinvesting capacity intentionally or whether it is dissipating into low-value activity. Pull attrition data by age bracket and role exposure. If your organization operates functions with significant AI tool adoption, check whether workers 55 and older are leaving at different rates than two years ago. If they are, investigate whether role changes, job insecurity, or capability mismatches are driving exits before attributing the trend to retirement timing alone. Review your engineering quality standards for the AI-era defect rate. If your teams use AI coding tools routinely, verify that code review protocols account for the higher defect rates now documented in independent pull request analysis. A review standard built for human-written code may be under-calibrated for the current output mix. If you don't own the final decision on any of these, bring the data to the person who does. The BCG survey, the Boston College research, and the CodeRabbit analysis each offer specific enough numbers to anchor a workforce planning or operations conversation. Framing these as structural gaps rather than AI adoption questions tends to move them out of the technology discussion and into the business-risk discussion where they belong. If your organization deployed AI tools to save time, what exactly did you plan to do with the time once it appeared? If you want to stay current on how AI is reshaping workforce conditions, career trajectories, and engineering quality across professional environments, and what those changes mean for the people living through them, Agenticism is where those stories live every day. For the curated weekly, monthly, and quarterly digest delivered to your inbox, subscribe at Agenticism on Substack. Sources BCG AI Workplace Research (via Forbes), View Article Center for Retirement Research at Boston College (via CNBC), View Article AI Code Quality Research, 2026 (via Tech Insider), View Article

  • July 14, 2026: The Three AI User Types, Only One Keeps Your Judgment Intact

    The professionals most at risk from AI over-reliance aren't the ones who use it least, they're the high-performers who use it constantly and have stopped noticing what they've quietly outsourced. A cluster of 2026 research studies has started to map this problem with unusual precision. The finding that should catch your attention: heavy AI use doesn't uniformly sharpen or dull thinking. It sorts users into three distinct behavioral patterns, and only one of those patterns preserves the independent reasoning skills that high-stakes decisions actually require. In this post. The Three User Clusters, what the 2026 research identified and where most experienced professionals land What "Balanced Support-Seeker" Actually Means in Practice, the specific habit that separates the cluster maintaining judgment from the ones that don't The Self-Diagnostic, a direct way to assess your own pattern this week without needing a study or a coach What to Do If You're in the Wrong Cluster, small, evidence-based adjustments that don't require abandoning the tools Most Experienced Professionals Are in the Cluster That Erodes Judgment A 2026 survey-based study by Bari et al., using machine learning to cluster AI users by behavioral pattern and then testing their unaided reasoning performance, identified three groups. The first group, over-reliant users, reaches for AI before attempting the problem independently. The habit feels efficient, and often is in the short term. But when the AI isn't available, or when the question is too nuanced for a reliable AI answer, their unaided performance is noticeably weaker than they expect it to be. The second group uses mixed strategies: sometimes independent, sometimes AI-assisted, with no consistent logic driving the choice. They haven't fully offloaded their thinking, but they also haven't built a deliberate practice that preserves it. The third group, which the research labels "balanced support-seekers," follows a consistent pattern: attempt the problem independently first, then use AI to pressure-test, verify, or extend. They treat AI as a check on their thinking rather than a replacement for it. According to Bari et al., this is the only cluster that maintained strong reflective problem-solving in unaided conditions. Complementary MIT research, which tracked 67 participants over four weeks on misinformation detection tasks, found that AI assistance improved immediate performance, but unassisted accuracy declined 15.3% by week four among participants who had been relying on AI assistance throughout. The reasoning capability weakens when it isn't regularly exercised. Michael Gerlich's 2025 research added a non-linear dimension: moderate AI use had minimal impact on critical thinking, but heavy reliance correlated with reduced critical thinking through cognitive offloading, the tendency to delegate mental processing to an external tool and stop performing that processing internally. A 2026 study published in Nature linked reliance on AI guidance to increased automation bias in decision-making. Automation bias is the documented tendency to accept outputs from an automated system at face value, including incorrect ones, because the system presented them with apparent confidence. Professionals who routinely accepted AI outputs were more likely to accept wrong answers when those answers sounded assured. The pattern across all four studies points in the same direction. The issue isn't whether you use AI, it's whether your current pattern keeps you in the practice of forming independent judgments. The Balanced Support-Seeker Habit Is Specific, Not Just a Mindset "Try first, use AI as backup" sounds obvious when written out. In practice, it requires a specific behavioral commitment that most busy professionals have quietly abandoned. Here is what it looks like in a working day. A finance professional reviewing a vendor analysis: the balanced support-seeker frames their own initial read, identifying what concerns them, what they'd want to dig into, what conclusion they'd tentatively draw, before opening the AI tool. Then they use AI to stress-test that read, surface what they might have missed, and challenge their assumptions. The over-reliant professional opens AI first, reads the summary, and works backwards from there. The conclusion still feels like theirs. The judgment call still feels independent. But the original framing, which is where most analytical errors get introduced, came from somewhere else. Action step. On your next significant analysis or decision, write two or three sentences of your own assessment before you open any AI tool. It takes 90 seconds and immediately tells you whether you have an independent view or whether you've been waiting for AI to give you one. The difference between the two approaches isn't speed or quality in the short term, it's what happens to your unaided reasoning over weeks and months. The MIT data suggests the decline starts within a four-week window, which means the habit compounds quickly in both directions. The Self-Diagnostic You Can Run This Week You don't need a survey or a structured benchmark to identify your own cluster. Three questions, answered honestly, are enough. 1. When you face an unfamiliar or complex question at work, what do you open first? 2. When AI gives you an answer, do you regularly form a competing view of your own before accepting it, or do you typically refine from what the AI has already given you? 3. If you had to make your three most consequential decisions from last month without AI assistance, how confident are you in what your unaided judgment would have produced? If your honest answer to question one is consistently "the AI tool," and question three produces hesitation rather than confidence, the Bari et al. research suggests you are likely in the over-reliant cluster. That's not a character flaw, it's an efficient habit that formed because the tools are genuinely good. Knowing where you stand is the starting point for changing the pattern. Most professionals who run this check discover they use mixed strategies: sometimes independent, sometimes AI-first, with no consistent logic. The mixed-strategy group isn't in immediate trouble, but they're also not building the deliberate practice that the balanced cluster maintains. What to Do If You're in the Wrong Cluster Small, consistent adjustments shift the pattern. The research doesn't suggest abandoning AI, it suggests changing the sequence. Build the "own view first" habit into your existing workflow. Before opening an AI tool on any decision that matters, write or say aloud your initial read. Even a rough one. This preserves the judgment-formation habit the balanced cluster maintains. It adds 60-90 seconds to your process. Reserve AI for verification and extension, not origination. Use it to find what you missed, not to tell you what to think. The distinction is small in terms of tool behavior and significant in terms of what happens to your reasoning over time. Deliberately practice unaided analysis on lower-stakes problems. The MIT four-week data showed decline even on relatively simple tasks. If every low-stakes question goes straight to AI, you're not preserving the capability for when you need it most. Reserve a category of routine judgment calls for unaided work. When AI gives you a confident-sounding answer on a consequential question, construct a dissenting view before accepting it. This is the direct antidote to automation bias, the documented pattern from the Nature 2026 study where AI-presented confidence increases acceptance of incorrect outputs. If you can't construct a plausible counterargument, that's a signal the AI may have closed your thinking before you've fully evaluated the question. Most professionals end up with a hybrid practice that reflects where each day's decisions fall on the stakes spectrum: AI-assisted for routine research and drafting, deliberately "own view first" for the decisions that define their professional judgment. That's not a compromise, it's the pattern the research identifies as sustainable. Try These Now Write your initial read before opening AI on the next complex question you face. Three sentences, rough and unpolished. Note whether you had an independent view or found yourself waiting for AI to frame the problem for you. Run the three self-diagnostic questions on your three most consequential decisions from last month. The confidence level on question three will tell you more about your current cluster than any formal assessment. Keep a five-day sequence log, "AI-first" or "own view first", on every analysis or decision where you reach for AI. No judgment required. You'll know your pattern by day three, and the data is genuinely useful. Pick one category of routine professional judgment and commit to handling it unassisted for four weeks. Not to prove a point, to keep the reasoning capability active and measurable. The MIT research suggests four weeks is a long enough window to detect whether the habit has already started to weaken. When did you last disagree with an AI answer on a question that actually mattered, and what did you do with that disagreement? The professionals who stay sharp in high-stakes decisions aren't the ones who use AI less, they're the ones who use it in a sequence that keeps their own judgment in the lead position. If you want to stay current on what AI means for individual professionals, the practical edge, not the organizational hype, Personal Agenticism is where those insights live. Subscribe at Agenticism on Substack for the curated weekly delivery. Sources Bari et al. 2026, AI User Clustering Study, View Article MIT ACM Study, Misinformation Detection and Unaided Accuracy Decline, View Article APA, AI Overreliance and Confidence, View Article Nature 2026, Cognitive Bias and AI Reliance, View Article Gerlich 2025, Cognitive Offloading and Critical Thinking, View Article

  • The AI Hardware Bottleneck Will Crack, and the Companies Built Around Scarcity Have the Most to Lose

    Lightmatter's photonic interconnect platform, which moves data between chips using light instead of electricity, hit a record 1.6 terabits per second per fiber in early 2026 and joined Nvidia's NVLink Fusion ecosystem in the same period. Cerebras reported that multiple financial services customers achieved 5 to 8 times faster model training cycles on its wafer-scale chips versus standard GPU clusters, with documented reductions in energy per training step. These are production deployments, arriving at the same moment that TSMC's advanced chip packaging capacity remains more than 80% allocated to Nvidia through 2026. The reason a VP of operations or finance should care right now is straightforward. The "we can never build enough" constraint that has driven AI infrastructure pricing for three years is beginning to loosen at the edges, and the companies whose pricing power depends on that scarcity are the ones most exposed. For organizations currently locked into expensive AI contracts, the negotiating equation is shifting. The Trend in Plain Sight Three distinct hardware shifts are converging, and each one chips away at a different part of the current cost structure. Optical interconnects are moving from prototype to production. Lightmatter's Passage L200 platform, unveiled in April 2025, offers co-packaged optics (a technology that integrates optical data transmission directly into the chip package, replacing power-hungry electrical connections) at 32 to 64 terabits per second. The company's liquid-cooled Laser NIC for higher rack densities shipped in May 2026. Critically, Lightmatter joined Nvidia's NVLink Fusion ecosystem, meaning optical interconnect technology is no longer positioned as an alternative to the dominant GPU architecture, it is becoming part of it. Ayar Labs shipped in-package optical prototypes to US customers for AI accelerators in February 2024, providing a second validation point that the integration path is advancing. Alternative chip architectures are capturing real enterprise spend. Groq's LPU chips (processors purpose-built for running AI models quickly, using fast on-chip memory instead of the expensive stacked memory modules that dominate current GPU designs) showed 10 times lower response latency per output token versus GPU baselines on Meta's Llama-3 70B model, according to Groq's public benchmarks. CoreWeave signed more than $1.6 billion in new GPU cloud contracts with enterprise AI teams in 2024, displacing portions of hyperscaler spend through lower per-GPU-hour pricing enabled by custom power infrastructure. These are not niche deployments, they represent documented enterprise budget moving away from the dominant providers. Cooling and power distribution are unlocking density that was previously impossible. Microsoft deployed immersion cooling at more than 100 kilowatts per rack in Azure AI regions in Q2 2024, then standardized direct-to-chip liquid cooling as part of its broader infrastructure push in 2025, and by 2026 was piloting in-chip microfluidics cooling (microscopic fluid channels built directly into the chip that remove heat up to three times more effectively than cold plates) in Phoenix and Mt. Pleasant. CoreWeave's liquid-cooled GPU clusters achieved 30% lower power usage effectiveness (the ratio of total facility power to IT equipment power, lower is better) than air-cooled equivalents as of June 2024. Financial services firms are moving fastest, driven by data residency rules and the need to protect proprietary risk model IP. Healthcare organizations face similar pressure from HIPAA requirements governing protected patient health information. Defense and government buyers are accelerating air-gapped specialized hardware deployments for sovereignty and security certification reasons. Why This Is Happening Three things changed in roughly the same 18-month window, and together they created conditions that did not exist in 2022 or 2023. 1. The packaging bottleneck became visible and quantified. TSMC's CoWoS capacity (the advanced chip packaging process that stacks high-bandwidth memory on top of AI processors) is targeted at 50,000 wafers per month and remains more than 80% allocated to Nvidia through 2026. When a single supplier controls the packaging process for the dominant AI chip and allocates nearly all of it to a single customer, every other chip designer faces a structural ceiling. That ceiling is now documented, not just rumored, and it is creating real incentive for alternatives. 2. The energy cost of the current architecture became a board-level problem. Power purchase agreements for new AI clusters face 18 to 24 month lead times in US regions. Lightmatter's photonic interconnect technology demonstrated sub-1 picojoule per bit energy consumption in inference workloads (running AI models to get answers from real business data), a meaningful reduction from electrical interconnect baselines. When energy costs are a constraint on how fast you can scale, a technology that cuts the energy per data transfer becomes a procurement conversation, not just an engineering one. 3. Open-weight AI models (models whose core inner workings are publicly shared, so companies can run them on their own systems without paying ongoing per-use fees) reached quality levels that made alternative hardware viable. Databricks reported that customers achieved 40% lower inference costs on open-weight models via optimized serving infrastructure versus proprietary API calls. When the model itself is free to run, the hardware economics dominate the decision, and that is exactly when architectural alternatives become competitive. It is like the shift from leasing specialized printing equipment at a premium to buying commodity printers once print volumes reached a threshold where ownership was cheaper. The equipment changed, but the real driver was the volume crossing a line where the per-unit cost of renting became indefensible. Key Numbers at a Glance 10x lower response latency per token, Groq LPU chips versus GPU baselines on Llama-3 70B, per Groq's public inference benchmarks (Groq, 2024) 5–8x faster model training cycles, Cerebras WSE-2/3 versus GPU clusters, reported by multiple financial services customers, with documented energy reductions per training step (Cerebras, April 2024) 40% lower inference costs, Databricks Mosaic AI customers running open-weight models on optimized infrastructure versus proprietary API calls (Databricks, 2024) 30% lower power usage effectiveness, CoreWeave liquid-cooled GPU clusters versus air-cooled equivalents (CoreWeave, June 2024) $1.6B+ in new contracts, CoreWeave enterprise GPU cloud signings in 2024, displacing portions of hyperscaler spend (CoreWeave, 2024) 80%+ of TSMC CoWoS capacity, allocated to Nvidia through 2026, constraining alternative chip designers and concentrating packaging supply (TSMC investor updates, 2024) 1.6 Tbps per fiber, Lightmatter record photonic interconnect bandwidth achieved in early 2026, with production-oriented co-packaged optics platforms now available (Lightmatter / OFC Conference, 2026) Here's Where This Points By 2027, specialized hardware providers and alternative cloud operators will capture a growing share of enterprise AI inference spend, potentially 15% or more of new budget in sectors where energy costs and data control matter most, if the documented cost savings from Groq, Cerebras, and CoreWeave continue to compound and open-weight model quality keeps closing the gap with proprietary alternatives. Photonic interconnects will reach limited commercial production in AI infrastructure by 2027, but broad deployment is more likely in the 2028 to 2029 window. The yield problems documented in 2023 and 2024 (integration yields below 70% in hyperscaler pilot tests) have not fully resolved. Lightmatter's NVLink Fusion integration and the Passage L200 platform suggest the technology is maturing, but the research notes Intel's volume availability on a similar timeline, and enterprise risk teams cite unproven long-term reliability in 24/7 training runs as a genuine adoption barrier. The "we can never build enough" constraint will ease for inference workloads before it eases for frontier model training. Running AI models to get answers from live business data is where alternative architectures (Groq's SRAM-centric design, Cerebras's wafer-scale approach) are already competitive. Training the largest next-generation models still requires the advanced packaging and memory bandwidth that TSMC and Nvidia dominate. These two workload types will diverge in their hardware economics over the next 24 to 36 months, and organizations that understand the distinction will make better procurement decisions. What This Means for the VP of Operations or Finance Leading AI Infrastructure Decisions If your organization is currently paying per-use fees for AI (charged based on how much you use it, like paying for electricity by the kilowatt-hour), the hardware shift matters to you in two ways. First, the cost floor for running AI is dropping, and it is dropping faster for high-volume, repetitive tasks than for complex reasoning work. Summarization, document classification, data extraction, and domain-specific generation are the workloads where Groq, Cerebras, and CoreWeave are already competitive. If your team is running these workloads through a major cloud provider's AI service at standard rates, you are likely paying a premium that will look increasingly hard to justify over the next 18 to 24+ months. Second, the energy and cooling constraints that have kept AI infrastructure pricing high are beginning to loosen, but unevenly. Microsoft's microfluidics cooling pilots and CoreWeave's liquid-cooled clusters are real efficiency gains, but power purchase agreement lead times of 18 to 24 months mean new capacity arrives slowly. The organizations that benefit first are the ones already in conversations with specialized providers, not the ones waiting for hyperscaler pricing to adjust on its own. For smaller teams without dedicated infrastructure budgets, the practical implication is different but equally concrete. Open-weight models running on platforms like Databricks Mosaic AI or Hugging Face Inference Endpoints are already 40% cheaper than proprietary API alternatives for the right workloads, per Databricks' reported customer data. The hardware efficiency gains flowing through specialized providers are what make those platforms economically sustainable at scale. Practical Next Steps In the next 30 days. Audit your current AI spend by workload type. Separate high-volume, repetitive tasks (summarization, classification, extraction) from complex reasoning tasks. The hardware economics for these two categories are diverging, and treating them as a single line item will lead to overpaying on one while underinvesting in the other. In the next 60 to 90 days. Run a cost comparison on one high-volume workload against a specialized provider. Groq's public inference benchmarks and CoreWeave's pricing are publicly available. Even if you do not migrate, having a credible alternative changes the negotiation, vendors know when you have options. For larger organizations. Ask your current cloud provider for a breakdown of what you are paying for AI services versus raw compute. The hyperscaler AI service markup (the additional fee that AWS, Azure, or Google Cloud charges on top of raw computing costs for their branded AI services) is the layer most exposed to competition from specialized providers. Knowing that number gives you a baseline for evaluating alternatives. For teams with energy or sustainability mandates. The cooling and power efficiency gains documented here (Microsoft's microfluidics work, CoreWeave's liquid cooling) are relevant beyond cost. Organizations with carbon commitments should be asking potential infrastructure partners for power usage effectiveness figures, not just per-GPU-hour pricing. The Second-Order Story The hardware efficiency story gets covered as an infrastructure narrative. The more consequential effect runs through the AI model providers and enterprise software companies that built their pricing on the assumption that scarcity would persist, which we have written about in prior posts. When an enterprise moves inference workloads to a Groq LPU cluster or a Cerebras wafer-scale deployment, it removes two fees simultaneously. The hyperscaler AI service markup and the model-provider per-use charge both disappear from the bill. The research documents this dynamic directly, Databricks customers achieving 40% cost reductions, CoreWeave displacing hyperscaler spend through custom power infrastructure. The revenue impact on OpenAI and Anthropic follows from the same migration math that affects Azure and AWS. Think of it like the shift from renting specialized lab equipment at a premium to buying commodity instruments once enough competitors entered the market. The rental company loses revenue, but so does the equipment manufacturer that had priced its products assuming the rental model would persist indefinitely. OpenAI and Anthropic are in the manufacturer's position here, not just the rental company's. Microsoft committed substantial capital to OpenAI with Azure OpenAI as the primary distribution vehicle. If Databricks and specialized providers are pulling inference workloads inside their own platforms on the same underlying infrastructure, Microsoft retains commodity compute revenue while losing the higher-margin AI services layer. Amazon invested $4 billion in Anthropic and positioned Claude on Bedrock as its premium AI offering. If Bedrock loses inference share to lower-cost specialized providers, that investment thesis faces pressure at exactly the moment it was expected to generate returns. The frontier model research funding loop is where the hardware efficiency story becomes structurally important. Training runs for the largest current AI models cost an estimated $50 to $100 million, and the next generation costs more. Both OpenAI and Anthropic fund these runs substantially from usage-based API revenue. If enterprise API revenue growth stalls on high-volume workloads, the predictable, large-contract customers that account for a disproportionate share of any usage-based business, the pace of frontier investment does not collapse immediately, but it becomes harder to sustain against Meta, which funds its AI research entirely from advertising revenue and has no equivalent API revenue exposure. The open-weight model releases that are enabling the hardware efficiency shift are being bankrolled by the only major AI lab with nothing to lose from lower inference costs. Enterprise software companies face problems too. Salesforce, SAP, and ServiceNow built AI upsell pricing on top of hyperscaler or proprietary model backend costs. If inference costs drop 40 to 80% through specialized hardware and open-weight models, the embedded AI premium across the enterprise software stack was priced into a world that is changing. The Salesforce and ServiceNow AI add-on licenses that drove the last two years of enterprise software revenue growth face a renegotiation they were not designed to absorb. What Could Slow This Down Photonic integration yields remain a genuine barrier. Hyperscaler pilot tests in 2023 and 2024 showed integration yields below 70%, delaying production deployment beyond 2026. Intel's optical I/O roadmap slips pushed volume availability to 2027. Lightmatter's progress is real, but the gap between a record benchmark and reliable 24/7 production deployment in a training cluster is significant, and enterprise risk teams are right to flag it. TSMC's packaging dominance is not dissolving quickly. CoWoS capacity remains more than 80% allocated to Nvidia through 2026. China's state-backed silicon photonics foundry lines at 300mm are announced but unverified at scale, and US export controls on advanced packaging equipment continue to limit photonic and 3D stacking supply chains for non-allied foundries through 2025. Supply diversification is a trajectory, not a current reality. Power infrastructure lead times offset hardware efficiency gains. New AI clusters face 18 to 24 month lead times for power purchase agreements in US regions. Microsoft's cooling innovations help with density inside existing facilities, but they do not accelerate the permitting and grid connection timelines that constrain new capacity. Organizations planning significant AI infrastructure expansion in 2025 and 2026 are working inside constraints that hardware efficiency alone cannot solve. Multi-year contracts and organizational inertia slow migration. Enterprise AI infrastructure decisions are not quarterly, they involve multi-year agreements, security review processes, and integration work that takes time. The economics may favor migration, but the switching costs are real, and the organizations most locked into existing hyperscaler agreements will move last regardless of the hardware signals. China's alternative hardware ecosystem has documented performance gaps. Huawei's Ascend clusters experienced 15 to 20% higher effective power draw than Nvidia equivalents in independent tests, attributed to software stack immaturity. The hardware roadmap is aggressive (600,000 Ascend 910C units targeted for 2026, doubling prior output), but software maturity typically lags hardware capability by 12 to 18 months in new architectures. Bottom Line By 2027, specialized inference providers and alternative hardware architectures will capture a meaningful share of enterprise AI workloads in sectors where energy costs, data control, and pricing flexibility matter most, with financial services and healthcare leading. The hyperscalers retain their advantages in frontier model training and managed tooling for complex workloads. The high-volume middle tier, where Groq, Cerebras, and CoreWeave are already competitive on documented benchmarks, is where the pricing power of the current dominant providers is most exposed. The companies that built their revenue models on the assumption that infrastructure scarcity would persist indefinitely are the ones with the most to rethink over the next 24 to 36 months. For your organization, the practical advantage is available now. Audit your workloads by type, run one cost comparison against a specialized provider, and enter your next contract renewal knowing what alternatives exist, or will exist relatively soon. Sources Lightmatter, Photonic chip energy benchmarks (1.6 Tbps per chip, sub-1 pJ/bit in inference workloads), May 2024. Foundational evidence for optical interconnect energy reduction path. https://www.ofcconference.org/news-media/exhibitor-news/lightmatter-achieves-record-1-6-tbps-per-fiber-to-accelerate-ai-optical-interconnect/ Lightmatter, Passage L200 co-packaged optics platform (32/64 Tbps versions), NVLink Fusion ecosystem integration, liquid-cooled Laser NIC, April 2025 / May–March 2026. Shows photonic interconnects moving from prototype toward production-oriented deployment within the dominant GPU ecosystem. https://lightmatter.co/; https://www.networkworld.com/article/3951672/lightmatter-launches-photonic-chips-to-eliminate-gpu-idle-time-in-enterprise-ai-data-centers.html Cerebras, WSE-3 customer throughput metrics: 4x training throughput per watt versus prior generation; multiple financial services customers reported 5–8x faster training cycles with documented energy reductions, April 2024. Validates wafer-scale architecture as a commercial alternative to GPU clusters for training workloads. Groq, LPU inference latency disclosures: 10x lower latency per token versus GPU baselines on Llama-3 70B, 2024. Commercial evidence for SRAM-centric design eliminating HBM interposers in inference clusters. CoreWeave, Enterprise contract announcements ($1.6B+ in 2024); liquid-cooled GPU clusters achieving 30% lower power usage effectiveness than air-cooled equivalents, June 2024. Documents specialized cloud operators capturing efficiency advantages and displacing portions of hyperscaler spend. Microsoft, Immersion cooling deployment reports (100+ kW rack density in Azure AI regions), Q2 2024; zero-water-evaporation cooling designs and direct-to-chip liquid cooling standardization, December 2024 / 2025; in-chip microfluidics cooling pilots in Phoenix and Mt. Pleasant, June 2026. https://www.microsoft.com/en-us/microsoft-cloud/blog/2024/12/09/sustainable-by-design-next-generation-datacenters-consume-zero-water-for-cooling/; https://news.microsoft.com/source/features/innovation/microfluidics-liquid-cooling-ai-chips/ TSMC, CoWoS capacity guidance: 50,000 wafers/month target for AI customers; 80%+ allocated to Nvidia through 2026, 2024 investor updates. Establishes the packaging bottleneck that is creating structural incentive for alternative architectures. Huawei, Ascend 910B optical switching fabric deployments in production training runs, Q4 2023; Ascend 910C/950 roadmap and 600,000-unit 2026 production target, September 2025. https://www.huawei.com/en/news/2025/9/hc-xu-keynote-speech; https://www.rcrwireless.com/20250930/ai-infrastructure/huawei-ai-chips-2 Intel, Memory architecture patent filings on interposer bypass, 2023; silicon photonics OCI chiplet and co-packaged optics IP portfolio with over 8 million photonic integrated circuits shipped historically, 2025–2026. https://www.intel.com/content/www/us/en/products/details/network-io/silicon-photonics.html Databricks, Mosaic AI customers achieved 40% lower inference costs on open-weight models via optimized serving infrastructure versus proprietary API calls, 2024. Key evidence for the economics driving workload migration away from proprietary model APIs. Ayar Labs, In-package optical I/O prototypes shipped to US customers for AI accelerators, February 2024. Second validation point for the optical interconnect integration path alongside Lightmatter. *Technical readers can find detailed customer metrics and benchmarks in the original announcements linked above.*

  • July 13, 2026: Zoom's Contact Center AI Hit 98% Containment. Only 23% of Enterprise Leaders Think Their Workforce Is Ready for That.

    In this post. Zoom reports a 98% chat containment rate and 25-point CSAT gain from AI virtual agents in its own contact center, along with broader survey data on enterprise adoption Hyster-Yale and NTT DATA embedded physical AI into manufacturing operations, cutting deployment timelines from months to weeks Kyndryl's 2026 People Readiness Report finds AI is in 57% of enterprise core processes, but only 23% of leaders believe their workforce can absorb it Two named organizations published specific AI deployment numbers this week. The Kyndryl research explains why those two are exceptions: 57% of enterprises have AI in core processes, but workforce readiness dropped six points year over year, and only 11% of organizations have hit both of their primary AI objectives. The deployment rate is climbing. The capability to extract results from those deployments is not keeping pace. Contact Center AI Is Posting Numbers. Zoom's Own Deployment Is the Lead Case. Zoom's analysis of AI virtual agents in contact centers draws on the company's own deployment and broader survey data. According to the company, its AI virtual agent achieved a 98% chat containment rate, meaning nearly all chat interactions were resolved without transfer to a human agent. Customer satisfaction scores moved from 55% to 80%, a 25-point gain. More than 1,000 agent hours were saved. Zoom also cites survey figures from its own research base: 69% of companies report AI improves customer experience, nearly three-quarters have achieved positive ROI, and 49% report revenue gains from agentic AI (systems that take sequences of actions to complete tasks, rather than just responding to a prompt). Because these figures come from Zoom's own customer base and survey outreach, they reflect organizations already using the tools and motivated to report positive outcomes. Independent corroboration would sharpen the picture. The 1,000-plus agent hours saved is the number that changes workdays for real people. Whether that capacity becomes redeployment into higher-value customer interactions or becomes the basis for headcount reduction is a decision made by leaders, not the technology. Contact center teams seeing these results in pilots should have a clear position on that question before the numbers start accumulating into a business case. Physical AI on the Factory Floor Is Compressing Deployment Timelines Hyster-Yale Materials Handling and NTT DATA announced a physical AI solution that embeds intelligence into manufacturing operations through sensor data, enabling real-time perception and action on the production floor. The system runs locally using edge computing, meaning processing happens on-site rather than through a remote cloud connection. For manufacturing environments where latency and connectivity reliability matter, that distinction is practical, not just technical. The stated result is deployment timelines cut from months to weeks versus legacy techniques. For manufacturing and operations leaders, that compression changes the economics of evaluation. Shorter deployment cycles reduce the cost of learning whether a system works in a real production environment, which makes it easier to run smaller experiments and build evidence before committing to scale. For frontline workers, embedded AI perception systems change what the day-to-day job looks like: less manual monitoring, more exception handling and oversight. Whether that transition means fewer operators or higher-value work per operator depends on how the organization manages the shift. The Hyster-Yale and NTT DATA announcement describes the technology and the timeline improvement. The workforce transition planning is the work that still sits with operational leaders. Kyndryl's Readiness Data Explains Why These Two Stories Are Outliers Zoom and Hyster-Yale stand out partly for the numbers themselves, and partly because those numbers are rare. Kyndryl's 2026 People Readiness Report, which surveyed 1,100 senior business and technology leaders across eight countries, captures why. 57% of enterprises now have AI embedded in core business processes, up from 35% a year ago. Despite that pace, only 23% of business leaders believe their workforce is fully prepared for the AI already deployed, a six-point drop from last year. Only 32% of organizations have achieved even one of their two primary AI objectives. Just 11% have hit both. Per Kyndryl's own research, 52% of leaders said finding employees with the right AI skills has become harder over the past year, and only one-third have fully implemented training programs designed to prepare staff to work alongside AI tools. 79% agreed that the pace of AI development will outpace their organization's ability to adapt its workforce, governance structures, and operating models. Kyndryl identifies a 9% cohort it calls "Pacesetters," organizations that are twice as likely to have fully implemented AI governance and 1.5 to 1.6 times more likely to report AI-driven revenue growth and improved innovation. The differentiator across that group is not which AI tools they chose. It's governance and workforce preparation running in parallel with deployment, rather than trailing it. The organizations publishing real AI results, whether on a contact center dashboard or a manufacturing floor, tend to be the same organizations that treated readiness as a parallel workstream, not an afterthought. The deployment rate across enterprise will keep rising. The readiness gap will be the deciding factor in who converts that deployment into outcomes, and who reports in next year's Kyndryl survey that they've still only hit one of two objectives. Act on These Now Map where your AI deployments are against where your workforce preparation programs are. If you have AI in production and fewer than one-third of affected staff have completed any structured preparation, you're in the majority of Kyndryl's sample, and also in the group that hasn't hit primary objectives. Define what "hours saved" means for your team before it defines itself. Whether freed capacity from AI automation becomes redeployment into higher-value work or becomes the basis for headcount decisions is a choice that should be made explicitly, not by default after the results accumulate. Identify your organization's governance readiness alongside its deployment readiness. Kyndryl's Pacesetter cohort separates from the rest on governance and workforce prep running concurrently with rollout, not after it. If your org has deployed without those in place, the sequencing gap is recoverable, but it has to be named first. Are the AI results your organization is presenting to leadership independently verifiable, or self-reported by the team and vendor running the deployment? The difference matters for trust in AI-driven business cases, and for accurately setting expectations about what comes next. If you want to stay current on how AI is reshaping contact center operations, manufacturing workflows, and enterprise workforce readiness, and what it means for the people and organizations living through it, Agenticism is where those stories live every day. For the curated weekly, monthly, and quarterly digest delivered to your inbox, subscribe at Agenticism on Substack. Sources Hyster-Yale / NTT DATA Press Release, View Article Zoom Blog. AI Virtual Agents in Contact Centers 2026, View Article Kyndryl / MarketScale. AI Adoption Up, Workforce Readiness Down, View Article

  • July 13, 2026: The BCG Coaching Experiment Should Change How You Think About Skill Gaps

    A BCG experiment with 139 employees found that one-to-one AI coaching taught a complex analytical skill, problem framing, 23% faster than a traditional virtual classroom. Among employees who were newer to the skill, the advantage grew to 32% larger gains than the workshop produced. That number comes from BCG Henderson Institute research published via HBR, comparing AI coaching against structured workshops of the kind most companies still rely on for professional development. It is a controlled comparison, not a vendor satisfaction survey. If you have a skill gap that matters for your next role or project and your company's L&D calendar won't address it for months, the data gives you a clear alternative path, with a specific design requirement attached. In this post. The BCG/HBR Numbers, what 139 BCG employees revealed about AI coaching speed, and what the novice-advantage tells you about where you stand in your development The Skills That Fit This Pattern, the capability types where AI coaching delivers, and why problem framing is a useful proxy for a whole category of judgment-heavy skills How to Design the Practice Loop Yourself, the three-part structure that converts AI coaching speed into durable capability rather than fast-fading surface knowledge Where Human Input Still Earns Its Place, the specific inflection points where skipping human accountability costs you more than any coaching fee saves Start Here, immediate actions calibrated to where you are in your skill development AI Coaching Delivered Real Speed Gains on a Genuinely Difficult Skill Problem framing is not a soft skill. It is the analytical ability to define what problem you are actually solving before you invest resources in solving it, one of the most consistently underestimated capabilities in any professional setting, and one that separates senior judgment from competent execution. BCG chose it deliberately. The experiment put 139 BCG employees through either one-to-one AI coaching or a traditional virtual classroom to build this skill. The AI group reached the same competency level 23% faster. Among employees newer to the skill, the gain was 32% larger than what the workshop produced. After a single session, 53% of participants rated the AI coach higher than human instruction, per BCG Henderson Institute's published findings. A separate Conference Board study found that 96% of users rated AI coaching as highly tailored to their needs, and 90% said they were comfortable with the experience. Among those tracking explicit career goals, 89% or more reported meaningful progress. The Conference Board's own research notes that these figures reflect users who opted into AI coaching, a group predisposed toward engagement, so the satisfaction numbers carry some selection lean. The BCG controlled experiment does not have that limitation: two groups, same skill target, measurable performance gap. The Skills That Fit This Pattern, and the Ones That Don't The BCG result generalises most cleanly to skills with a learnable structure, capabilities that have identifiable components, common failure patterns, and observable practice opportunities. Problem framing fits well. So do negotiation preparation, structured communication, performance feedback delivery, and many elements of career navigation. These are skills where AI coaching does something a workshop cannot. It responds to your specific scenario rather than a generic case study. It gives you immediate feedback on the draft negotiation approach you are preparing for a real conversation next Tuesday. It asks the follow-up question a thoughtful human coach would ask, and does so at 11pm when you are actually thinking through the problem. The pattern that works: a skill with learnable components, a real upcoming application, and enough complexity that you benefit from repeated practice with feedback rather than a single read-through. The pattern that fits less cleanly: skills that are fundamentally relational at their core. Building organisational trust over time, navigating a specific political dynamic inside your company, reading a room in ways that require deep contextual knowledge of particular people. AI coaching can help you prepare for those situations. It cannot replace the accumulated judgment that comes from actually living through them with someone who knows the territory. How to Design the Practice Loop Yourself Speed is only useful if it converts to capability that holds under pressure. The failure mode in AI coaching, flagged in the Conference Board research, is that the interaction feels productive while it is happening but does not get consolidated into durable behaviour without deliberate design on your part. The loop that works has three components: 1. Identify one specific skill with a real deadline. Not "improve my executive presence." Something with edges: "Be able to frame the strategic problem in three sentences before I present to the leadership team on August 14." The AI coach needs a real target to be useful. Vague goals produce polished but generic practice. 2. Run compressed practice sessions with your own scenarios. Use the AI to work through your actual cases, not generic examples. If you are practising problem framing, bring the real project you are working on. Ask the AI to challenge your framing, offer an alternative definition, and push back on your assumptions. Twenty minutes with real material beats three sessions with abstract exercises. 3. Schedule one human review at the midpoint and one at application. A peer, a mentor, or a manager who knows your context. This is not about validation, it is about catching the blind spots the AI cannot see because it does not know your organisation's history, your specific stakeholders, or the implicit constraints on what "good" looks like in your environment. Human check-ins also create accountability that sustains practice past the first session. Action step. Before your next AI coaching session, write one sentence naming the specific skill, the real application, and the date you need it. If you cannot write that sentence, spend the first session on goal clarification rather than skill practice, that is itself a productive use of the time. Where Human Input Still Earns Its Place The BCG experiment showed AI coaching outperforming workshops on skill acquisition speed. It did not show AI replacing the full range of what a skilled human coach delivers. Human coaches add distinct value at three specific inflection points. First, when the gap is primarily perceptual rather than technical. If you cannot yet see why your current behaviour is a problem, if the gap is invisible to you, a human who knows your context will surface it faster and more accurately than an AI working from your self-description alone. Second, when the stakes involve real professional consequence. A promotion conversation, a difficult performance review you need to deliver, a negotiation where the relationship will outlast the outcome. AI can help you prepare. The accountability for how it lands belongs to a human who shares professional context with you. Third, when you need someone to challenge the frame you are working inside, not just the execution. AI coaching is highly responsive within the frame you give it. A human coach who knows your history will challenge the frame itself, and that is often where the highest-value insight sits. The practical result is a hybrid approach rather than a substitution. Use AI for velocity: the repeated, low-friction practice sessions that build the underlying skill. Use human input for calibration: the moments when you need someone who knows both the skill and your specific situation to confirm that what you have built actually fits where you are heading. Most professionals still default to waiting for the next workshop or finding budget for a coach. The BCG data shows you can start building the underlying capability this week, at your own pace, on your own schedule, as long as you design the accountability structure yourself rather than assuming the AI will provide it. Start Here Pick one skill gap with a real deadline in the next 60 days and write a single sentence defining it specifically enough that a colleague could understand it without a follow-up question. If you cannot write that sentence in under a minute, your first AI session should focus on goal clarification. Bring your actual work into the practice session. The BCG gains came from contextualised coaching, not generic exercises. Use a real negotiation you are preparing for, a real presentation you are building, or a real feedback conversation you have been avoiding. Schedule two human check-ins before you start, one at the midpoint of your practice period and one the week before the real application. Put them on the calendar now, before the AI sessions begin. Without them, the accountability structure that converts fast learning into durable behaviour does not exist. If you are newer to a skill, the BCG data suggests your gains will be larger. The 32% novice advantage in the BCG experiment is meaningful. This approach is especially productive if you are building a capability that is genuinely new for you rather than refining one you already have solid foundations in. After each AI coaching session, ask yourself. Could you teach this skill component to a colleague right now? If the answer is no, the session produced understanding, not capability. Go back and work through one more real application before moving on. If you want to stay current on what AI means for individual professionals, practical development leverage, evidence-based tools, and clear signals on what actually works versus what sounds plausible at a conference, Personal Agenticism is where those insights live. Subscribe at Agenticism on Substack for the curated weekly delivery. Sources BCG Henderson Institute, How Gen AI Could Transform Learning and Development, View Article Conference Board, AI Can Provide Career Coaching, But Humans Still Matter, View Article Dr. Philippa Hardman, HBR Study Summary via LinkedIn, View Article

  • May 5, 2026: Pentagon Contracts, Wall Street Deals, and 73,000 Layoffs Later — AI Has a New Job Title

    This weekend's news cycle wasn't short on concrete moves: military deployment agreements, billion-dollar enterprise JVs, pre-launch government safety reviews, and a wave of layoffs framed explicitly around AI. Taken individually, each story has a clear business angle. Taken together, they mark a shift from AI as something companies are experimenting with to something they're committing institutional capital and policy authority to. Here's what happened. Seven AI Companies Get Pentagon Clearance for Classified Networks The Defense Department announced on May 1 that Google, Microsoft, Amazon Web Services, Nvidia, OpenAI, SpaceX, and Reflection AI have reached agreements to deploy their AI systems on classified military networks "for lawful operational use." The Pentagon's stated goal is to build what it called "an AI-first fighting force" with "decision superiority across all domains of warfare." A further expansion to include Oracle was confirmed by May 4. Notably absent from the list: Anthropic. The company has been in a legal standoff with the Defense Department since late 2025, refusing to lower its safety guardrails for autonomous weapons and mass surveillance use. Anthropic won an injunction in March blocking the Pentagon's attempt to brand it a supply-chain risk, and it remains excluded from this round of classified network agreements. The split is worth tracking closely. The companies that signed on agreed to military use cases that Anthropic explicitly declined. For enterprises evaluating AI vendors, this isn't just competitive positioning, it's a values-based fork in the road that will shape what these models are optimized for and what constraints get relaxed over time. Knowing which vendor made which choice matters when you're deciding what to embed in your own operations. Anthropic's Mythos Model Pushes the Government Into Pre-Launch Oversight The classified network deals weren't the only government AI action that week. On Tuesday, the Commerce Department's Center for AI Standards and Innovation (CAISI) announced that Google, Microsoft, and xAI have agreed to provide federal agencies pre-launch access to evaluate new frontier models before public release. That now brings all five major labs, including OpenAI and Anthropic from prior agreements, into a voluntary pre-release evaluation program. The catalyst was Anthropic’s Claude Mythos Preview, released in April. The model demonstrated the ability to autonomously discover and exploit thousands of previously unknown zero-day vulnerabilities across major operating systems, web browsers, and government infrastructure, at a speed and scale no human red team could match. Anthropic limited access to 11 partner organizations and the UK’s AI Security Institute, citing the risks of broader release. The UK reported that Mythos uncovered thousands of unpatched vulnerabilities. That capability reportedly accelerated discussions in Washington. The resulting voluntary evaluation framework still lacks statutory authority and is supported by a small team of under 200 staff, which is currently America’s closest equivalent to formal AI oversight. Editor’s Note (May 8, 2026): This section was updated to reflect the EU AI Act high-risk compliance deadline postponement, which was agreed on May 7, 2026, shortly after the original publication date of this article. How much this changes before the EU AI Act’s high-risk obligations take effect (now December 2027, delayed from the original August 2026 date) remains an open question the industry is actively pricing in. OpenAI and Anthropic Both Launched Enterprise JVs on the Same Day On May 4, TechCrunch confirmed that OpenAI and Anthropic each launched separate enterprise AI joint ventures on the same day, backed by Wall Street, designed to embed their models directly inside large companies. Anthropic's venture is backed by Blackstone, Permira, and Hellman & Friedman. OpenAI's, internally called DeployCo, formally "The Development Company", is raising $4 billion at a $10 billion pre-money valuation from 19 investors, including TPG, Bain Capital, Brookfield, and Advent International. OpenAI is committing up to $1.5 billion directly. Both ventures are using the forward-deployed engineer model: embed engineers inside client teams, gain preferred sales access through PE portfolio companies, and accelerate enterprise adoption faster than traditional channel partnerships allow. Combined estimated value: approximately $11.5 billion. Reuters reported by May 5 that both JVs are already in acquisition talks, looking to purchase AI services firms to build out deployment capacity quickly. The practical implication for enterprise buyers is immediate. If your company sits in a PE portfolio, expect your AI vendor relationship to arrive pre-packaged with your PE firm's preferred provider. If you're running an AI platform evaluation, the path to a real deployment contract now runs through investment relationships, not just procurement cycles. Palantir Reports 85% Revenue Growth, The Fastest Since Its IPO On May 4, Palantir reported Q1 2026 earnings: 85% revenue growth, the company's fastest expansion since it went public in 2020. Palantir builds AI-powered data analysis and decision-making platforms across government and commercial clients, and has been one of the clearest beneficiaries of enterprises shifting from AI pilots to production deployment. The 85% number is significant beyond Palantir itself. The company's commercial revenue has been growing alongside its government contracts, which means defense adoption and enterprise adoption are moving in parallel, not the staggered sequence most analysts projected two years ago. If you're building ROI models for AI deployment at your organization, Palantir's earnings curve is one of the cleanest third-party data points available on what full-scale AI integration actually looks like financially. Coinbase Cuts 14% of Its Workforce and Redesigns Around AI Agents On Tuesday, Coinbase CEO Brian Armstrong announced roughly 700 job cuts , 14% of the global workforce, and framed it explicitly as an AI-driven restructuring, not just a crypto market response. "AI is bringing a profound shift in how companies operate, and we're reshaping Coinbase to lead in this new era," Armstrong wrote. The cuts come with a full operational redesign: management layers are being reduced to a maximum of five below the CEO, and the company is creating what Armstrong called "AI-native pods" , potentially one-person teams directing AI agents that collectively handle the work previously requiring engineers, designers, and product managers together. His description of the new model is one of the blunter articulations yet of where this is heading: "We are not just reducing headcount and cutting costs, we're fundamentally changing how we operate: rebuilding Coinbase as an intelligence, with humans around the edge aligning it." That framing, AI at the center, humans in a supervisory and alignment role at the perimeter, is increasingly common in executive memos. The question worth asking is whether your team is building the skills to be in that supervisory layer, or the ones that sit inside it waiting to be directed. Tech Sector Layoffs Cross 73,000 in 2026, With AI as the Stated Driver Coinbase is the most recent, but far from the largest. As of this week, more than 73,000 roles have been eliminated across 95 tech companies in 2026, according to data from Layoffs.fyi. Amazon cut 30,000 corporate and tech jobs since October, which is roughly 10% of its corporate workforce. Oracle cut thousands, explicitly tied to ramping AI infrastructure spending. Dell reduced headcount by 10% for the third consecutive year. The pattern is consistent: companies are investing aggressively in AI infrastructure while simultaneously shrinking the teams AI is positioned to replace. Amazon and Oracle aren't doing this because AI tools aren't working. They're doing it because, in their operational assessment, the tools are working well enough to justify accelerating the transition ahead of stabilizing the workforce. For teams still in planning mode on AI deployment, the most useful reframe is no longer "what could AI do here?" It's "which of these roles is on the 18-month restructuring shortlist?" That distinction separates organizations that get ahead of this shift from those that find out about it from HR. Australia Moves Toward AI Enforcement, and the EU Is Three Months Out Outside the US, Australia's financial and data protection regulators threatened enforcement proceedings this week against companies demonstrating inadequate AI controls , which is a concrete shift from the advisory posture most regulators have held for the past two years. Australia isn't an outlier. The EU AI Act's full enforcement window opens in December 2027. For organizations operating in or serving EU markets, the requirements are operational, not aspirational: agent identity management, comprehensive audit logs, documented human oversight protocols, and the ability to revoke an AI's operating access within seconds. Nominal human involvement, a human technically in the loop, is no longer sufficient. Regulators have made clear they want to see humans who can actually understand how AI makes decisions and override them. The enforcement window is the point at which governance slide decks stop being sufficient. For US enterprises with EU exposure, that deadline is already inside the planning horizon for most IT cycles. The Coinbase restructuring is the story that carries the week's clearest implication forward. What Armstrong described, which is a company rebuilt as an intelligence, with humans at the edges aligning it, is the operational model that every enterprise JV, every Pentagon contract, and every pre-launch government evaluation is ultimately pointing toward. The question isn't whether that model arrives. It's whether your organization is building the capability to operate inside it or waiting to inherit the outcome. If you want to stay ahead at the intersection of AI, automation, and human performance, where technology meets psychology, processes, and real workplace behavior, subscribe to Agenticism. We cut through the hype to deliver practical insights for leaders focused on making people, processes, and technology work better together.

  • Private AI Is No Longer a Compliance Workaround, It's Becoming the Default Infrastructure for Regulated Industries

    CoreWeave secured multi-year contracts with two US hedge funds for isolated AI inference clusters, running AI models to process live trading and risk data entirely within dedicated, controlled environments, and delivered 40% lower per-use costs than public APIs while satisfying data residency requirements. That happened in 2024. By mid-2026, CoreWeave had expanded to 43 active data centers covering more than 850 megawatts of power, acquired Core Scientific for $9 billion to lock in 1.3 gigawatts of additional capacity, and announced a $6 billion Pennsylvania data center investment alongside European expansions explicitly designed to meet GDPR and jurisdictional data rules. If you lead a team in financial services, healthcare, defense contracting, or any regulated sector, the infrastructure decision your organization makes in the next 18 months will determine whether AI becomes a controlled, auditable business asset or a persistent compliance liability. The question is no longer whether private AI hosting is viable. The question is whether your organization is moving fast enough to capture the cost and control advantages before your competitors lock in the best infrastructure contracts. The Trend in Plain Sight Financial services moved first, and the numbers explain why. A major US bank deployed Databricks Mosaic AI private endpoints, meaning AI models running entirely within the bank's own dedicated computing environment, with no data leaving the bank's systems, and achieved zero external data transfers for risk model inference over nine months. CoreWeave's hedge fund contracts show the same pattern: isolated clusters, 40% cost reduction versus public pay-per-use services, and documented data residency compliance. These organizations are not running experiments. They are running live production workloads, meaning actual day-to-day business processes with real data and real users, inside controlled infrastructure. Healthcare followed on patient data protection grounds. A healthcare provider deployed Snowflake Cortex private AI functions, AI capabilities running entirely inside the provider's existing Snowflake data environment, with no information sent to external model providers, and passed an internal audit with no protected patient health information (PHI, governed by HIPAA's strict privacy and security rules) leaving the system. The provider did not need to build new infrastructure. It used the data platform it already owned. That pattern, running AI inside existing data systems rather than sending data out to external AI services, is becoming the default architecture for regulated healthcare organizations. Defense and government drove the hardware-level shift. Groq delivered on-premises inference hardware, physical computing equipment installed inside a defense contractor's own facility, achieving sub-10 millisecond response times on classified workloads with no cloud connectivity required. Groq's GroqRack program, which provides plug-and-play rack configurations for air-gapped environments (systems with no external network connections), raised $650 million in June 2026 and now operates 13 global data centers alongside its on-premises offering. NVIDIA's DGX Cloud Lepton marketplace, launched in May 2025, connects organizations to partner GPU capacity from CoreWeave and Lambda Labs with explicit support for region-specific data residency and hybrid deployments. The pattern across all three verticals is the same. Regulated organizations are moving AI inference, the everyday "using" phase of AI, as opposed to the initial training phase, inside their own controlled perimeters. The drivers are cost, compliance, and control, in that order, and all three are now pointing the same direction. Why This Is Happening Now Three things changed between 2023 and 2026 that did not exist together before. Open-weight AI models reached enterprise-grade quality on routine tasks. Open-weight models are AI models whose core inner workings are publicly shared, so organizations can run them on their own systems without paying ongoing per-use fees to the original creator. Meta's Llama series, and similar models from Mistral and others, now perform comparably to proprietary models on high-volume, repetitive tasks like document summarization, classification, data extraction, and domain-specific generation. The quality gap that once justified paying premium per-use prices has closed on those workloads. It has not closed on complex, novel, multi-step reasoning, that still favors frontier proprietary models, but the routine work that makes up the majority of enterprise AI volume no longer requires it. The cost math flipped at scale. Think of it like deciding whether to keep renting specialized equipment every time you need it, or bringing the work in-house once volume makes ownership cheaper and safer. At low volumes, renting wins: no capital commitment, no maintenance. At high volumes, ownership wins: the per-unit cost drops below the rental rate, and you control the asset. Enterprise AI crossed that threshold. Lambda Labs reported 3x GPU utilization gains for enterprise customers running private fine-tuned models, models trained further on a company's own specific data to perform better on that company's tasks, versus shared hyperscaler instances. CoreWeave's hedge fund contracts demonstrate 40% cost reduction at production scale. Regulatory pressure shifted from optional to mandatory. The EU AI Act reached full applicability in May 2026, with a deadline of December 2027. US financial services regulators have intensified data residency requirements. HIPAA enforcement on AI-processed patient data has sharpened. Multiple 2026 analyses from NTT DATA, Cloudera, and others identify sovereign AI, AI infrastructure operating under a specific jurisdiction's legal and technical controls, as a strategic necessity for regulated industries, not a premium option. Organizations that previously treated private hosting as a compliance workaround are now treating it as the baseline architecture for any AI workload touching sensitive data. Key Numbers at a Glance 40% lower per-use cost, CoreWeave's isolated inference clusters versus public API pricing for two US hedge fund clients, while meeting data residency requirements. (CoreWeave, 2024) Zero external data transfers, Databricks Mosaic AI private deployment for a major US bank running risk model inference over nine months, confirmed by the bank's internal compliance review. (Databricks, 2024) 3x GPU utilization improvement, Lambda Labs enterprise customers running private fine-tuned models versus shared hyperscaler instances, per the company's reported case studies. (Lambda Labs, 2024) $9 billion, CoreWeave's acquisition of Core Scientific in 2026, securing 1.3 gigawatts of power capacity to support dedicated AI infrastructure at scale. (CoreWeave, 2026) $650 million, Groq's June 2026 fundraise to scale its AI inference cloud and on-premises GroqRack program for regulated and air-gapped environments. (Groq, 2026) 13 global data centers, Groq's operational footprint as of mid-2026, alongside active on-premises hardware deployments for defense and regulated industry clients. (Groq, 2026) Here's Where This Points Current migration patterns and the documented cost, compliance, and quality signals make three trajectories increasingly likely over the next 24 to 36 months. High-volume, routine AI workloads in regulated industries will largely move to private or sovereign infrastructure by 2028. Financial services, healthcare, and defense are already past the pilot stage on this. The combination of regulatory mandates, documented cost savings, and open-weight model quality on routine tasks creates conditions where contract renewals will favor private stacks. Organizations that have not begun this migration by 2027 will face a compressed timeline when their current API contracts come up for renewal. The hyperscalers, AWS (Amazon), Azure (Microsoft), and Google Cloud, will retain compute revenue but lose the higher-margin AI services layer on these workloads. AWS's expansion of Bedrock PrivateLink support in February 2026 and Microsoft Azure's sovereign cloud options are defensive moves, not growth strategies. They keep the underlying computing revenue while ceding the premium AI service fees to specialized providers and data platform incumbents like Databricks and Snowflake. If the migration patterns observed so far continue, specialized providers could capture 15 to 20 percent of new enterprise AI infrastructure spend by 2028, concentrated in regulated verticals. Complex, frontier-level AI tasks will remain on proprietary models for the foreseeable future. The quality gap on multi-step reasoning, novel analysis, and tasks requiring the most capable models has not closed. OpenAI and Anthropic retain defensible positioning on those workloads. The migration pressure falls on the high-volume middle, the repetitive, data-intensive work that built the enterprise AI revenue story of the past three years, not on the frontier reasoning tasks where performance differences still justify proprietary pricing. What This Means for the VP of Compliance and Legal Operations If you sit at the intersection of AI adoption and regulatory accountability, the infrastructure shift described above is simultaneously your biggest risk management opportunity and your most pressing organizational challenge. Private and sovereign AI infrastructure gives your organization something public API deployments cannot deliver: a documented, auditable chain of custody for every piece of sensitive data that touches an AI model. When a regulator asks where your patient data went during AI processing, "it stayed inside our Snowflake account" is a categorically stronger answer than "it was processed by an external model provider under their terms of service." The Snowflake Cortex deployment that passed a healthcare provider's internal audit with zero PHI egress is the kind of evidence that changes how compliance conversations go. The challenge is organizational, not technical. Your legal and compliance team needs to be involved in AI infrastructure decisions before they are made, not after. The pattern of enterprises reverting to public APIs after failed private deployments, one financial services firm abandoned an air-gapped setup after four months due to model update complexity, almost always traces to compliance requirements being added after the architecture was chosen. The organizations succeeding at private AI deployment started with the compliance requirements and worked backward to the infrastructure, not the other way around. For smaller legal and compliance teams without dedicated AI infrastructure resources, the practical path is through existing data platforms. If your organization already uses Snowflake or Databricks, both now offer AI capabilities that run entirely inside your existing environment without external model calls. You do not need to build a private cloud. You need to understand what your current data platform can do with the AI features it already has. Practical Next Steps In the next 30 days. Audit where your organization's AI workloads currently send data. For every AI tool your teams use, map whether the data stays inside your systems or leaves to an external model provider. Most organizations discover they have more external data exposure than their compliance documentation reflects. That audit is the starting point for every conversation that follows. In the next 60 days. If you use Snowflake or Databricks, schedule a technical review of their current private AI capabilities. Both platforms now offer in-perimeter AI functions as generally available features, not experimental add-ons. The question is whether your current data governance setup can support them. This is a procurement and architecture conversation, not a research project. In the next 90 days. For high-volume, repetitive AI workloads, document processing, classification, extraction, summarization, run a cost comparison between your current per-use API spend and what a dedicated or private deployment would cost at your actual volume. The math changes significantly above certain thresholds, and most organizations have not done this calculation recently. Even if you do not migrate, having a credible alternative changes the negotiation. Vendors know when you have options. For smaller teams. The on-premises hardware path (Groq's GroqRack, Lambda Labs' Echelon clusters) carries significant upfront capital cost and is not realistic for most mid-size organizations. The more accessible path is data-platform-native AI, running models inside Snowflake or Databricks rather than calling external APIs. The compliance benefits are comparable; the capital requirements are not. The Second-Order Story The enterprise infrastructure shift gets the attention. The downstream consequences for the AI industry's economics deserve equal scrutiny. When an organization moves production AI inference inside its own Snowflake or Databricks environment using an open-weight model, it removes two fees simultaneously. The hyperscaler AI service markup (the additional fee that AWS, Azure, or Google Cloud charges on top of raw computing costs for their branded AI services like Bedrock, Azure OpenAI, or Vertex AI) disappears, and so does the model provider's per-use charge. A mid-size financial services firm running $5 million annually in OpenAI API calls on high-volume document processing can reduce that spend by $3 to $4 million by moving to a fine-tuned open-weight model inside its existing data platform. The migration pays back in under a year at that scale. The enterprises doing this math are not edge cases, they represent the predictable, high-volume API customers that account for a disproportionate share of revenue at any usage-based AI business. Think of it like what happened to long-distance telephone revenue in the early 2000s. The per-minute charges that built the business model collapsed not because the calls stopped, but because the underlying infrastructure became cheap enough that the premium layer lost its justification. OpenAI and Anthropic built their enterprise revenue models on a world where running AI at scale required their infrastructure. Open-weight models running on dedicated clusters are the equivalent of internet calling, same outcome, different economics, and the premium disappears. The investor theses behind the major model providers face a version of this pressure that has not yet surfaced in public financials. Microsoft committed over $10 billion to OpenAI across multiple tranches, with Azure OpenAI as the primary distribution vehicle for that investment's returns. If Databricks and Snowflake are pulling inference workloads inside their own environments on the same Azure infrastructure, Microsoft retains commodity compute revenue while losing the higher-margin AI services layer it funded through the OpenAI relationship. Amazon invested $4 billion in Anthropic and positioned Claude on Bedrock as its premium AI offering. If Bedrock loses inference volume to private data-platform deployments, Amazon's investment thesis weakens at exactly the moment its strategic AI bet does. Both hyperscalers added open-weight models to their managed services in 2024 and 2025, a defensive measure that acknowledges the pressure without resolving it. The frontier R&D funding loop is where this becomes structurally important for the industry. Training runs for the most capable AI models cost an estimated $50 to $100 million per run, with each generation costing more. OpenAI and Anthropic fund these runs substantially from usage-based API revenue. If enterprise API revenue growth stalls on high-volume workloads, the tier most exposed to private hosting migration, the pace of frontier investment does not collapse immediately, but it faces sustained pressure against a competitor, Meta, that funds its AI research entirely from advertising revenue and has no usage-based revenue to protect. Meta's open-weight release strategy is disrupting OpenAI and Anthropic's enterprise revenue model while Meta itself faces no equivalent disruption. That asymmetry compounds over time. The enterprise software incumbents face a quieter version of the same problem. Salesforce, SAP, and ServiceNow built AI upsell pricing on the assumption that inference costs would remain at a level that justified their embedded AI premiums. A Salesforce Einstein license priced on $0.01-per-token inference looks different when enterprises can run comparable models at $0.001 per token inside their own infrastructure. The AI copilot premium across the enterprise software stack was priced into a world where inference costs stayed high. CoreWeave and Fireworks.ai pricing already shows that floor dropping for dedicated deployments. If it continues, the upsell logic that drove enterprise AI software revenue growth over the past two years faces a renegotiation it was not designed to absorb. What Could Slow This Down Several forces are actively working against the migration timeline described above. Model update complexity is the most underreported barrier. Multiple enterprise pilots of fully air-gapped inference failed because keeping models current inside isolated environments requires ongoing engineering work that most organizations underestimated. One financial services firm reverted to Bedrock after four months specifically because of patching and update complexity. Air-gapped deployments trade compliance simplicity for operational complexity. Organizations that succeed at this invest in the internal capability before they migrate, not after. Open-weight model performance gaps persist on specialized tasks. Three consulting firms abandoned private hosting trials in 2024 because open-weight models did not perform adequately on their specific use cases. The quality gap has closed on routine, high-volume tasks. It has not closed on complex reasoning, nuanced judgment, or highly specialized domain work. Organizations that migrate the wrong workloads will revert. The discipline of distinguishing which tasks belong on private infrastructure and which still require frontier proprietary models is not yet common inside most organizations. Capital barriers remain significant for mid-size organizations. GPU supply constraints and 12 to 18 month lead times for dedicated clusters raise the capital commitment well above what pay-as-you-go APIs require. Two mid-size manufacturers canceled sovereign cloud contracts before deployment due to upfront hardware reservation costs. CoreWeave's $9 billion Core Scientific acquisition addresses this at the infrastructure level, but the capacity it unlocks takes time to reach enterprise customers as available, affordable dedicated clusters. Regulatory fragmentation adds compliance overhead. US state-level data localization rules remain inconsistent across jurisdictions, which means a private hosting architecture that satisfies California's requirements may not satisfy a different state's rules. Organizations operating across multiple US jurisdictions face compliance overhead that partially offsets the compliance benefits of private hosting. This is a solvable problem, but it requires legal and compliance involvement from the start of the infrastructure design process. Hyperscaler volume discounts are narrowing the cost gap. Model providers continue offering volume discounts that reduce the cost advantage of private open-weight deployments for organizations that have not yet reached the scale where ownership economics clearly win. For organizations running moderate AI volumes, the public API option remains cost-competitive in the near term. The migration economics become compelling at higher volumes and longer time horizons. Bottom Line By 2028, private and sovereign AI infrastructure will be the default architecture for regulated industry AI workloads in financial services, healthcare, and defense, with the migration concentrated on high-volume, repetitive tasks where open-weight model quality and dedicated cluster economics have already made the case. The hyperscalers keep the underlying compute revenue. OpenAI and Anthropic keep the complex reasoning workloads where frontier model performance still justifies proprietary pricing. The high-volume middle, the document processing, classification, extraction, and domain-specific generation that built the enterprise AI revenue story of the past three years, is actively in play, and the organizations that priced their businesses on that middle holding are the ones with the most to rethink. For your organization, the audit and the math are where the advantage is built. Map where your AI workloads send data today, run the cost comparison at your actual volume, and understand what your existing data platforms can already do with their native AI features. You do not need to build a private cloud to capture most of the compliance and cost benefits. You need to know what you already own and what it can do. Sources CoreWeave, Enterprise contract announcements for isolated inference clusters, data center expansions, European sovereign regions, Core Scientific acquisition. Documented 40% cost reduction versus public APIs for hedge fund clients; 43 active data centers and 850+ MW power by end-2025; $9B Core Scientific acquisition securing 1.3 GW capacity. (2024–2026) https://www.coreweave.com/ai-data-centers Databricks, Mosaic AI private deployment documentation. Zero external data transfers confirmed for a major US bank's risk model inference over nine months. (2024) Groq, GroqRack on-premises program details, $650M fundraise, 13 global data centers. Sub-10ms latency on classified defense workloads with no cloud connectivity; plug-and-play rack configurations for air-gapped environments. (2024–2026) https://groq.com/newsroom/groq-raises-usd650m-to-scale-its-ai-inference-cloud-business Snowflake, Cortex AI Functions and Agents general availability release notes. In-perimeter LLM inference with no external model calls; healthcare provider audit confirmation with zero PHI egress. (2024–2026) https://docs.snowflake.com/en/release-notes/2025/other/2025-06-02-cortex-aisql-public-preview Lambda Labs, GPU cloud utilization case studies and private supercluster announcements. 3x GPU utilization improvement for enterprise customers running private fine-tuned models versus shared hyperscaler instances; SOC 2 Type II compliance for private tenancy. (2024–2026) https://lambda.ai/ NVIDIA, DGX Cloud Lepton marketplace announcement, May 2025. Connects developers to partner GPU capacity (CoreWeave, Lambda) with explicit region-specific sovereignty and data residency compliance support. (May 2025) https://nvidianews.nvidia.com/news/nvidia-announces-dgx-cloud-lepton-to-connect-developers-to-nvidias-global-compute-ecosystem AWS, Bedrock PrivateLink/VPC endpoint expansion. Extended private VPC-only access to additional endpoints including distributed inference support. (February 2026) https://aws.amazon.com/about-aws/whats-new/2026/02/amazon-bedrock-expands-aws-privatelink-support-openai-api-endpoints/ Microsoft Azure, Azure AI Studio sovereign cloud expansion for US government-adjacent workloads. Directional signal toward regulatory-driven infrastructure segmentation; no volume metrics disclosed. (Q1 2024) Hugging Face, Inference Endpoints enterprise tier with dedicated hardware reservations and private networking. Lowers the barrier for enterprises running models outside public endpoints. (Q4 2023–2024) Red Hat + CoreWeave/Azure, Red Hat AI Inference (llm-d orchestration) validated on CoreWeave CKS and Azure AKS. Open orchestration layer reducing operational complexity of private inference management. (May 2026) https://www.redhat.com/en/blog/red-hat-ai-inference-brings-llm-d-any-managed-kubernetes-starting-coreweave-and-microsoft-azure Cloudera / NTT DATA, 2026 sovereign AI analyses identifying private and perimeter-bound AI deployments as strategic necessity for regulated industries amid EU AI Act full applicability. Directional signal; no proprietary metrics cited. (2026) https://www.cloudera.com/resources/faqs/sovereign-ai.html Groq / NVIDIA licensing, NVIDIA licensed Groq LPU technology in a reported ~$20B deal (December 2025), with Groq remaining independent and continuing on-premises and inference cloud offerings. (December 2025–June 2026) https://techcrunch.com/2026/06/22/ai-chipmaker-groq-confirms-650m-raise-re-staffs-after-nvidias-20b-not-acqui-hire-deal/ Technical readers can find detailed customer metrics and benchmarks in the original announcements linked above.

  • July 10, 2026: S&P Global Rebuilt Its Operating Model. HCLTech Just Signed $1.14B to Manage One for a Fortune 50 Company.

    In this post: S&P Global redesigned its Market Intelligence operating model around agentic AI, with executive leadership changes attached HCLTech signed a $1.14 billion, 5-year contract to run a Fortune Global 50 company's global digital workplace and enterprise networks on an AI-driven model SCWorx deployed an AI-assisted data management system for healthcare supply chain attributes, rolling out to selected customers now Anthropic launched a finance-specific agent marketplace connecting pre-built workflows to Moody's, Morningstar, PitchBook, and other major data sources A $1.14 billion contract to hand over digital workplace and network operations to an AI-driven model is a specific kind of commitment. It is not a procurement decision or a pilot extension. It is an operations handoff at Fortune 50 scale, and it arrived the same week S&P Global announced it was restructuring its core Market Intelligence operating model around the same logic. S&P Global Reorganized Market Intelligence, Not Just Its Tools On July 6, S&P Global announced a new operating model for its Market Intelligence division. The stated goal: align with customer needs in an AI-driven market by pairing S&P's data and domain expertise with integrated AI-powered tools, workflows, and experiences. The announcement included executive leadership changes alongside the structural redesign, which signals this is an organizational commitment rather than a product update. For a company whose product is information and analysis delivered to financial professionals, the distinction matters. S&P isn't adding AI features to existing workflows. It is restructuring the operating model around where AI can be the delivery mechanism, an approach it describes as accelerating agentic solutions and platform capabilities. The announcement does not detail implementation timelines or measurable outcome targets, which is typical at this stage. Operating model redesigns at this scale usually take 12-18 months before showing up in customer experience or financial metrics. If you work with S&P Global Market Intelligence products or compete in the financial data space, the trajectory here is visible even without the numbers. The HCLTech Contract Describes What an AI Operating Model Actually Is On July 3, HCLTech announced a $1.14 billion strategic contract with a Europe-headquartered Fortune Global 50 company. The contract runs from July 2026 through December 2031, with an option to extend for five more years. HCLTech shares rose more than 7% on the announcement. The scope is precise: establish an AI-driven operating model to transform and manage the client's global digital workplace and enterprise networks. In practice, this means AI handles device provisioning, software deployment, access management, incident routing, and service requests across what is presumably hundreds of thousands of employees worldwide. Humans set policy and handle exceptions. AI executes the operations. An AI operating model, as described in the analysis of this deal, is different from an AI tool. It is the organizational and technical infrastructure that runs your business using AI, not as a productivity booster for employees, but as the primary mechanism for delivering services. For enterprise network operations, it means AI monitors, diagnoses, and in many cases auto-remediates connectivity, performance, and security issues across global infrastructure. Network operations centers that once employed large teams of engineers shift to smaller teams managing AI systems rather than managing the infrastructure directly. For the employees inside this Fortune Global 50 company, this restructuring means their IT support, device management, and network reliability will be driven by AI systems. The announcement is clean on commercial terms. What it does not address is the transition experience for the operations teams being reorganized around it. SCWorx Targets the Data Foundation Under Healthcare Supply Chain AI On July 8, SCWorx Corp. announced the deployment of its AI-assisted Data Management Model for healthcare supply chain product attributes. The system automates classification, normalization, enrichment, and governance of supply chain data , which is the layer that determines whether a hospital's procurement system knows what it is actually buying and from whom. The deployment is currently rolling out to selected customer engagements, with broader availability planned. SCWorx has not disclosed specific outcome metrics or customer names at this stage. Healthcare supply chain data is notoriously difficult. Products listed under multiple names, inconsistent classification systems across vendors and group purchasing organizations, and manual governance that is expensive and error-prone. The AI-assisted approach targets this foundation layer specifically. The practical implication for supply chain and procurement professionals in healthcare is that data standardization has a high potential for return on investment. If your classification and normalization layer is messy, the automation built on top of it will be too. SCWorx's deployment addresses this problem, though results will vary considerably depending on the data quality each customer brings to an implementation. As always, conduct your own research before buying. Vendors Are Building Purpose-Specific Financial AI, With a Compliance Deadline Approaching The developments above involve named organizations restructuring operations. On the vendor side, Anthropic launched a dedicated financial services agent marketplace with pre-built templates for banks, asset managers, hedge funds, and insurers. These integrate Moody’s, Morningstar, PitchBook, Verisk, and SS&C Intralinks data directly into Claude workflows for credit analysis, underwriting, deal diligence, and market abuse monitoring. Claude Opus 4.7 leads the Vals AI Finance Agent benchmark. This is the practical direction. Vendors are delivering industry-specific tools instead of raw models. Under the EU AI Act, high-risk systems, including credit scoring, AML transaction monitoring, and insurance underwriting, require documented conformity assessments. Fines can reach €30 million (~$34 million). Today, the compliance burden sits squarely with the deploying institution, not the vendor. Anthropic’s marketplace helps with capabilities, but it does not satisfy regulatory responsibility. As we noted in a prior piece on the Workday lawsuit, it will be interesting to see if that clean split holds once vendor liability cases play out. Operating Models Are Being Rebuilt, Not Just Upgraded The S&P Global and HCLTech stories share structural logic. Both represent organizations treating AI not as a capability their people use, but as the primary mechanism through which operations are delivered. The $1.14 billion price tag on the HCLTech contract is one indicator. The five-year duration is another. You do not sign a five-year operations handoff for something you plan to reverse. The human dimension these announcements tend to compress is the experience of the people inside the organizations being restructured. Operating model redesigns produce cleaner unit economics on paper. They also produce a sustained period of uncertainty for operations teams whose roles shift from executing work to governing the AI systems that execute it. Act on These Now Map which of your operations match the pattern HCLTech is delivering. High-volume, rule-based functions where AI can execute and humans handle exceptions. IT help desk, procurement data management, document classification are natural candidates to start with. Knowing which functions in your organization fit this model is the prerequisite for any informed conversation about whether to build, buy, or outsource. Audit your supply chain data quality before the AI layer goes in. SCWorx's deployment targets classification and normalization specifically because bad data produces bad automation. If your organization is evaluating AI-assisted procurement or inventory decisions, the data foundation is the important place to start the assessment, not the vendor demo. If your AI systems touch credit decisions, AML monitoring, or insurance underwriting and you operate in the EU, the high-risk compliance deadline is now December 2027 (delayed from the original August 2026 date). This is not optional, so take action now. Vendor marketplace launches do not satisfy conformity assessment requirements. Your institution owns that responsibility. Vendor marketplace launches do not satisfy conformity assessment requirements. Your institution owns that responsibility, and the timeline is now weeks, not months. Before any AI operating model contract reaches the signature stage, define what "humans handle exceptions" means in practice. In a global enterprise, the exception rate and the escalation structure matter as much as the automation scope. Push for specifics on both before the ink dries. If you contribute to or execute within an operations function being evaluated for this kind of redesign, start documenting your domain expertise now. The roles that survive operating model shifts tend to belong to people who can articulate what the AI gets wrong, not just what it does. That knowledge is valuable, but only if it is visible. If you want to stay current on how AI is changing enterprise operating models and what it means for the people and organizations living through it, Agenticism is where you will find weekly operating model breakdowns that deliver clear and practical perspectives. For the curated weekly, monthly, and quarterly digest delivered to your inbox, subscribe at Agenticism on Substack. Sources S&P Global PR Newswire — View Article SCWorx GlobeNewswire — View Article HCLTech AI Operating Model — The DAILY Brief — View Article Anthropic Finance Agent Marketplace — RepresentAI — View Article

  • July 10, 2026: Your Prompts Didn't Get Worse, the Game Changed Around Them

    The reason your prompts stopped getting better results isn't that you need cleverer wording. Models have absorbed the standard techniques, role assignment ("act as an expert strategist"), dramatic persona instructions, and elaborate opening setups, so thoroughly into their training that these approaches no longer differentiate anyone's output. What differentiates results in 2026 is what surrounds the prompt, not the prompt itself. In this post. The Technique Plateau, why basic prompting tricks have stopped working and what practitioners report is actually happening inside modern models Context Engineering Replaces Clever Phrasing, the specific shift high-output professionals are making, from prompt wording to surrounding architecture Constraints and Guardrails Are the New Leverage, how explicitly defining what you don't want produces more reliable results than obsessing over what you do want Build Your Self-Evaluation Loop, a simple, zero-setup method for getting the model to audit its own output against your real success criteria Five Changes You Can Make Today, specific adjustments you can apply in your existing tools right now Basic Prompt Techniques Have Been Absorbed Into the Models Themselves Practitioners who work closely with leading AI models, Claude, ChatGPT, Gemini, reported in mid-2026 that the gimmicky phrasing era is functionally over. Role-play openers, dramatic persona assignments, and "think like a genius" instructions were useful when models needed those cues to activate certain reasoning patterns. According to discussions in active practitioner communities, those patterns are now default behaviors. The model doesn't need you to tell it to think carefully. It already does. What this means practically: if your prompting approach is largely unchanged from 2024, you are running yesterday's playbook on a different game. The outputs aren't bad. They're just generic. Good-enough-to-use, but not calibrated to your specific situation, constraints, or professional judgment. The ceiling isn't the model's capability. It's the absence of structure around it. Context Engineering Replaces Clever Phrasing, and Requires No Technical Background Practitioners and prompt researchers have converged on a term for the emerging practice: context engineering. It means deliberately designing what goes into a session before you type your first question, not after you're disappointed with the answer. The shift in practice looks like this. Instead of spending ten minutes crafting a perfect prompt for a one-off session, you spend that time once building a standing context document, a short, plain-text note you paste or upload at the start of any substantive AI session. It tells the model who you are, what project you're working on, what your actual role and constraints are, and what a good output looks like in your professional context. Action step. Build one standing context document this week. Three to four paragraphs covering your professional role, the project or domain you're working in most often, the audience for your typical outputs, and two or three things a good result always includes. Paste it at the start of your next five AI sessions and compare the output quality against your last five without it. A senior finance professional, for example, might write: "I'm preparing analysis for a board-level audience that has no patience for caveats without data. Good output here means a clear recommendation in the first sentence, three supporting data points, and explicit acknowledgment of the two most likely objections." That context, provided once, produces more reliable output than twenty iterations of prompt wording. This approach works inside every major browser-based AI tool you already use. If you have Google Workspace through your employer, you already have Gemini with contractual data protection, meaning Google does not use your work content to train its public models. The standing context document works there exactly as it does in Claude or ChatGPT. Constraints Outperform Aspirations in Prompt Design Most prompts focus on what you want. The higher-leverage move, according to practitioners studying model behavior in 2026, is specifying what you explicitly do not want. Leading AI models are trained to be helpful, which means they default toward comprehensive, balanced, diplomatically hedged output. Without explicit constraints, that is exactly what you get, thorough, unoffensive, and often not quite right for a high-stakes professional situation. Explicit constraints change this. They're not complicated. They look like: "Do not include caveats about data limitations unless a specific data gap materially changes the recommendation." "Do not summarize what I already told you, start directly with the analysis." "Do not suggest additional research. Work with what I've provided." "If you are uncertain, say so in one sentence and move on. Do not hedge every paragraph." Practitioners describe this as defining failure modes before the session starts, rather than correcting them after. The model's default helpfulness becomes an asset rather than a liability when it has clear boundaries on what "helpful" means in your specific context. Action step. In your next high-stakes AI session, preparing a presentation, drafting a recommendation, analyzing a complex situation, write three "do not" constraints before you write your main question. Notice whether you spend less time editing the output. Build Your Self-Evaluation Loop in Two Sentences The most underused technique in professional AI use costs nothing and requires no new tools. After you receive any output you're going to act on, send one follow-up message: "Grade this response against the criteria I gave you. Be direct about where it falls short and rewrite those sections." This works because modern models can evaluate their own outputs against explicit criteria more reliably than they can produce a perfect output in a single pass. You're adding a second pass that catches the gaps your own review might miss when you're pressed for time. The prerequisite is that you actually stated success criteria somewhere in the session, which is why the context document and constraints described above matter. Without criteria, there's nothing to evaluate against. With them, the model can flag where it hedged when you needed directness, where it listed options when you needed a recommendation, and where it used language your audience won't understand. Most professionals find in the first week of using this pattern that the second pass adds more value than re-prompting from scratch, and it takes about thirty seconds. Action step. Use this exact two-sentence follow-up in your next three substantive AI sessions and track whether you spend more or less time editing the final output. The Professionals Pulling Ahead Have Shifted From Incantation to Architecture The distinction that matters is this: clever prompting is an input optimization. Context engineering, constraint design, and evaluation loops are system design. Input optimization has diminishing returns because you're competing with every other professional trying to find slightly better wording for the same model. System design compounds over time because the structures you build, standing context, constraint libraries, self-evaluation habits, get better as you refine them. One practitioner framing from mid-2026 captures it well: the valuable work now is "structured problem design, clear constraints, guardrails, validation loops." Not the incantation. The architecture. For a non-technical senior professional, this is good news. None of these techniques require understanding how AI models work internally. They require understanding your own work, your audience, your constraints, your definition of a good outcome. That is exactly what experienced professionals already know about their domain. The move is applying that knowledge to how you set up every AI session, not to how you word the prompt. The professional who builds a strong standing context document, writes explicit constraints, and uses a two-sentence self-evaluation loop will consistently out-produce the one who spent the same time searching for the perfect opening line, because the former is building a system, and the latter is still looking for a magic spell. Five Changes You Can Make Today Build a standing context document this week. Three to four paragraphs covering your role, current project focus, typical audience, and what good output looks like in your professional context. Paste it at the start of every substantive session for two weeks and notice what changes. Add three "do not" constraints to your next high-stakes prompt. Before you write what you want, write what you explicitly do not want, hedging, unnecessary caveats, options instead of recommendations, summaries of what you already told the model. Compare the first draft you receive against your last five attempts without constraints. Use the two-sentence self-evaluation follow-up once per day. After receiving an output you'll act on, send: "Grade this against the criteria I gave you. Be direct about where it falls short and rewrite those sections." Track whether your editing time drops over two weeks. Stop re-prompting from scratch when you're disappointed. Re-prompting from scratch signals the model got something wrong without telling it what. Instead, name the specific failure. "The recommendation was buried in paragraph four, I need it in the first sentence" is far more effective than restating the original question with slightly different wording. Ask yourself this. If a new junior analyst joined your team today, could you hand them a one-page document explaining your role, your standards, your audience, and your definition of a good deliverable? If you couldn't, you haven't given your AI that foundation either, and that gap is why the outputs feel generic. If you want to stay current on what AI means for individual professionals, not the organizational hype, but the practical edge, Personal Agenticism is where those insights live. Subscribe at Agenticism on Substack for the curated weekly delivery. Sources Reddit, Prompt Engineering Is Dead in 2026, View Article Digital Applied, Advanced Prompt Engineering Techniques 2026, View Article The AI Corner, ChatGPT and Claude Power User Setup Guide 2026, View Article

  • Your AI Agents Are Already Acting Like Insiders. Most Organizations Haven't Noticed.

    Gartner projects the average Fortune 500 company will run more than 150,000 AI agents by 2028, up from fewer than 15 today. Only 13% of organizations believe their governance is adequate to manage that scale, according to Gartner's April 2026 guidance on agent sprawl. The gap between those two numbers is where your delegation of authority framework quietly became a liability. The core problem is structural, not technical. Delegation of authority matrices, the formal documents that define who in your organization can approve what, were designed around humans. A person has a job title, a manager, a defined scope, and an access review every 90 days. When that person changes roles or leaves, IT revokes their credentials. AI agents, the persistent software programs that now execute multi-step workflows across your CRM, HR system, email, and financial platforms, do not fit that model. They accumulate access, retain memory of prior interactions across sessions, and continue operating long after the conditions that justified their original permissions have changed. They behave, in practice, like high-privilege insiders. They just never appear on anyone's insider threat watchlist. The Trend in Plain Sight The pattern shows up across every major enterprise platform released in the past two years. Microsoft's Copilot Studio agents retain memory across sessions and execute multi-step workflows with persistent tool access, without requiring re-authorization for each action. As of May 2026, Microsoft has added a centralized governance layer called Agent 365 that includes agent inventory, permission controls, and memory management. The fact that Microsoft had to build that governance layer after the agents were already in production at 180-plus enterprise tenants tells you something about the sequence in which these capabilities arrived. Salesforce Agentforce deployments grant agents read and write access to CRM records, with memory of prior customer interactions carried forward across sessions. ServiceNow's Vancouver release, from late 2023, introduced agents that retain workflow context and execute across IT, HR, and security modules simultaneously. Google's Vertex AI Agent Builder enables stateful agents, meaning agents that remember what they have done and use that history to inform future actions, with access spanning multiple Google Workspace and Cloud resources. The Deloitte 2026 State of AI in the Enterprise survey of 3,235 leaders found that 23% of firms are already using agentic AI, with that figure projected to reach 74% within two years. Only about 20% of those firms have mature governance models in place. Deployment is running roughly three to four years ahead of governance, and the gap is widening. Regulated industries are moving fastest, which creates an irony. Financial services and healthcare firms face the strictest data rules, so they have the most incentive to deploy agents that can handle sensitive workflows efficiently. They also have the most to lose when those agents accumulate privileges that no compliance officer formally approved. Why This Is Happening Now Three things changed in roughly 18 months that made this problem qualitatively different from earlier automation risks. Agents gained persistent memory. Earlier robotic process automation tools, the software robots that have been running in back offices for years, executed defined scripts and stopped. They did not remember. Current AI agents, built on large language models, maintain context across sessions. An agent that helped a customer resolve a billing dispute in January still carries that interaction history in March. That memory shapes how the agent behaves, what it accesses, and what it infers it is permitted to do. Memory is a form of accumulated privilege, and most identity management systems have no mechanism to audit or expire it. Agents gained broad tool access. Anthropic's Claude computer-use API, released in October 2024, allows agents to operate a browser and file system directly. Microsoft, Salesforce, and ServiceNow agents connect to multiple enterprise systems simultaneously. The scope of a single agent's reach now routinely exceeds what any individual human employee would be granted in a standard access review. Machine identities already vastly outnumber human ones. CyberArk's 2025 Identity Security Landscape report found that machine identities, the digital credentials assigned to software systems, outnumber human identities by more than 80 to one. AI agents are the fastest-growing category within that figure. Your identity management team was already stretched before agents with persistent memory entered the picture. Your delegation of authority matrix is like a building's key card system. It tracks which humans can enter which rooms, and it revokes access when someone leaves. Now imagine a cleaning crew that never clocks out, learns the building layout over months, and gradually starts opening doors it was never formally issued a key for, because no one programmed the system to notice that a cleaning crew is different from an employee. That is roughly the situation with persistent agents and current access controls. Key Numbers at a Glance 150,000+ agents per Fortune 500 company, Gartner's projected average by 2028, up from fewer than 15 today; only 13% of organizations report adequate governance (Gartner, April 2026) 80. 1 ratio, machine identities already outnumber human identities at that ratio, with AI agents the fastest-growing category (CyberArk, 2025 Identity Security Landscape) 74% of enterprises, Deloitte's projected share deploying agentic AI within two years; only ~20% have mature governance in place today (Deloitte, 2026) 41% reduction in unauthorized access events, CyberArk's reported outcome after deploying identity controls for 12,000 non-human agent identities at two Fortune 100 financial firms (CyberArk, 2025) 700+ organizations exposed, Okta's documented case of the Salesloft Drift breach, where long-lived digital access keys, called OAuth tokens, outlasted their intended purpose and exposed connected organizations (Okta, November 2025) 30-day privilege expiration, ServiceNow's Vancouver agent deployments at three healthcare systems include automatic access expiration after 30 days of inactivity, one of the few documented examples of time-bounded agent authority in production (ServiceNow, 2023) Here's Where This Points Current deployment rates and the documented gap between agent capabilities and governance frameworks point toward three developments over the next 24 to 36 months. Agent identity will become a formal compliance category. The EU AI Act's Article 14, with enforcement beginning December 2027, requires proof of authorization at the time an agent executes an action, not just at the time the access was originally granted. That is a materially different standard from how most enterprise access controls work today. US frameworks are moving in the same direction: NIST's AI Risk Management Framework is developing an agentic profile, and the Cloud Security Alliance published governance standards for non-human agent identity in May 2026. Organizations that treat agent governance as an IT housekeeping task are likely to find it reclassified as a compliance requirement before 2027. Delegation of authority matrices will need agent-specific sections. The current approach, where agents inherit human user permissions or operate under generic service accounts, creates an audit trail that satisfies neither regulators nor internal risk teams. The emerging architectural standard, documented by Okta, Orchid Security, and Strata in 2025 and 2026, treats agents as distinct identity classes with explicit delegation chains, short-lived credentials that renew rather than persist indefinitely, and memory treated as an auditable privileged action. Organizations that build this architecture now will have a structural advantage when regulators formalize the requirement. The insider threat framework will expand to cover non-human actors. Palo Alto Networks' Unit 42 incident response team formally classified autonomous AI agents as a new insider threat category in 2025 and 2026, based on documented cases of agents retaining access after the human employees who authorized them changed roles or left. The tools that detect insider threats today, the behavioral analytics platforms that flag unusual data access patterns, were trained on human behavior. Agents move faster, access more systems simultaneously, and generate activity patterns that current tools flag as false positives rather than genuine risks. That detection gap will close, but the organizations that close it proactively will avoid the incidents that force everyone else to close it reactively. What This Means for Chief Compliance Officers and General Counsel If you are in company deploying AI agents, the central question your team needs to answer is not "what can our agents do?" It is "who formally authorized them to do it, and can we prove that authorization is still valid?" Your current delegation of authority matrix almost certainly does not have an answer to that question. The matrix was built to track humans. Agents are not humans. They do not appear in org charts, they do not have performance reviews, and they do not trigger the offboarding workflow when a project ends. But they do access sensitive data, execute financial transactions, modify records, and in some cases communicate with customers on your organization's behalf. The productivity case for agents is documented and growing. CyberArk's deployment at two Fortune 100 financial firms reduced unauthorized access events by 41% after implementing proper identity controls, which means the firms were running with those unauthorized access events before the controls were in place. The firms that are moving carefully on governance are not slowing down their agent deployments; they are making those deployments auditable, which is what allows them to scale further without accumulating compliance exposure. For smaller organizations without a dedicated identity security team, the immediate practical question is simpler. Do you know which agents are currently running in your environment, what systems they can access, and whether any of them were authorized by an employee who has since changed roles? If you cannot answer all three, you have an inventory problem before you have a governance problem. Practical Next Steps In the next 30 days, run an agent inventory. Most organizations deploying Microsoft Copilot, Salesforce Agentforce, or ServiceNow agents do not have a complete list of which agents are active, what permissions they hold, or when those permissions were last reviewed. Microsoft's Agent 365 governance platform, released in May 2026, provides this for Copilot Studio agents. If you are on other platforms, the inventory may require manual work. Do it anyway. You cannot govern what you have not counted. In the next 60 days, map your three highest-privilege agents to your delegation of authority matrix. Pick the agents with the broadest system access and ask who formally authorized this scope. Is that person still in the same role? Has the business purpose for the agent changed? This exercise will surface gaps faster than any vendor tool. In the next 90 days, establish a minimum standard for new agent deployments. Before any new agent goes into production, require three things: a named human owner who is accountable for the agent's actions, a defined expiration or review date for its permissions, and a documented list of the systems it can access. ServiceNow's 30-day inactivity expiration at three healthcare systems is a workable model for organizations that need a starting point. For larger organizations with existing IAM infrastructure, the vendors with the most mature agent-specific controls as of mid-2026 are Okta (AI Agent Lifecycle Management), CyberArk (non-human identity controls), and Palo Alto Networks (Prisma AIRS for agent discovery and runtime protection). None of these are complete solutions yet, but each addresses a distinct part of the problem. For smaller teams, the most practical near-term approach is simpler: treat every agent deployment like a new employee hire. Define the scope before you deploy, not after. The Second-Order Story The governance gap creates a downstream problem that goes beyond any individual organization's compliance posture. The enterprise software vendors whose revenue depends on broad agent adoption have a structural incentive to ship capabilities faster than governance frameworks can absorb them. Microsoft, Salesforce, and ServiceNow all built agent memory and broad tool access before they built the governance controls. Microsoft's Agent 365 platform arrived roughly two years after Copilot Studio's persistent memory features. Salesforce's Agentforce launched with CRM read/write access before formal delegation controls were available. The sequence is not accidental; it reflects competitive pressure to ship features. But it means every enterprise customer is running a governance deficit that the vendor created and is now selling solutions to close. Think of it like a contractor who installs a swimming pool without a fence, then returns six months later to sell you the fence. The pool is genuinely useful. The fence is genuinely necessary. But the customer is paying twice, once for the capability and once for the control, and the gap between installation and fencing is when the liability accumulates. The identity security vendors, CyberArk, Okta, Palo Alto Networks, and a newer category of agent-specific governance tools including Orchid Security and Aembit, are positioned to capture meaningful spend from this gap. CyberArk's deployment of controls for 12,000 non-human identities at two Fortune 100 firms is an early indicator of the contract sizes available. If Gartner's 150,000-agent projection holds, the addressable market for agent identity governance is substantially larger than the current privileged access management market, which was already a multi-billion dollar category. The organizations most exposed to forced spending are the ones that deployed agents aggressively in 2024 and 2025 without updating their governance frameworks. They will face a choice between retrofitting controls under regulatory pressure, which is expensive and disruptive, or accepting the compliance exposure, which is increasingly untenable as EU AI Act enforcement begins and US frameworks follow. The cost of retrofitting legacy identity and access management platforms to track persistent agent state is estimated to be a multi-year project at large banks, according to the foundational research signals. Organizations that start the inventory and framework work now are buying time against that cost. The frontier AI model providers, OpenAI and Anthropic, face a quieter version of this pressure. If enterprises respond to governance concerns by limiting agent scope and reducing the volume of actions agents take, overall API call volumes decline. An agent that requires explicit re-authorization for each sensitive action makes fewer autonomous calls than one operating with persistent broad access. The governance trend and the cost-reduction trend point in the same direction: toward lower per-enterprise API consumption on high-volume agentic workloads. That is not the growth trajectory either company's revenue model was built on. What Could Slow This Down The governance gap will not close quickly, and several forces will keep it open longer than the urgency of the problem would suggest. Existing identity and access management platforms were not built for agents. Mapping agent memory state to static role definitions is a technical problem that most enterprise IAM systems cannot currently solve. The retrofitting cost at large organizations is measured in years, not quarters. Compliance teams do not have the headcount to review agent memory logs at scale. The volume of activity generated by even a modest agent deployment exceeds what human reviewers can process using current tools. The behavioral analytics platforms designed to detect insider threats generate excessive false positives when applied to agent activity, because agents move faster and access more systems simultaneously than the models were trained to expect. Organizations that slow agent deployment to mature governance frameworks first will, in the short term, deploy fewer agents than competitors who move faster. That pressure is documented in the research: pilot programs at two consulting firms abandoned broad agent access after auditors flagged the lack of formal delegation updates. The firms that paused were doing the right thing. They also fell behind on deployment timelines. US regulatory frameworks are still catching up. SEC and FTC guidance continues to reference human decision-makers as the primary accountability unit. The EU AI Act's execution-time authorization requirement is the clearest regulatory forcing function currently in effect, but its direct reach in US markets is limited to organizations with EU operations or customers. Bottom Line By 2027, agent identity governance will be a formal compliance requirement for any organization operating in regulated industries or with EU market exposure, not an IT best practice. The organizations that treat the current window as an opportunity to build agent inventory, delegation frameworks, and access controls will enter that regulatory environment with documented processes. The ones that do not will be building those processes under audit pressure, at higher cost, with less time. Your agents are already acting with insider-level access. The question is whether your governance framework treats them that way. The tools to close that gap exist now. The cost of waiting is compounding every quarter that agent deployments scale ahead of the controls designed to manage them. Sources Microsoft, Copilot Studio governance and memory features, Agent 365 GA (May 2026). Centralized agent inventory, permission controls, lifecycle oversight, and memory management via Dataverse; safe-sharing detection for credential oversharing. https://www.microsoft.com/en-us/microsoft-copilot/blog/copilot-studio/new-and-improved-agent-governance-intelligent-workflows-and-connected-app-experiences/ Gartner, AI agent sprawl management guidance (April 2026). Projects Fortune 500 average will reach 150,000+ agents by 2028 from fewer than 15 today; only 13% of organizations report adequate governance. https://www.gartner.com/en/newsroom/press-releases/2026-04-28-gartner-identifies-six-steps-to-manage-artificial-intelligence-agent-sprawl Okta, AI agent authorization drift and lifecycle management (November 2025). Documents "authorization drift" where digital access keys outlive their intended purpose; cites Salesloft Drift breach exposing 700+ organizations via long-lived tokens; introduces purpose-built agent identity framework with instant revocation. https://www.okta.com/blog/ai/ai-agent-security-when-authorization-outlives-intent/ CyberArk, 2025 Identity Security Landscape report (April 2025). Machine identities outnumber humans by more than 80 to 1; AI agents identified as distinct high-privilege insider threat vectors; documents 41% reduction in unauthorized access events after deploying controls for 12,000 non-human agent identities at two Fortune 100 financial firms. https://www.cyberark.com/press/machine-identities-outnumber-humans-by-more-than-80-to-1-new-report-exposes-the-exponential-threats-of-fragmented-identity-security/ Deloitte, State of AI in the Enterprise 2026 (3,235 leaders surveyed). Agentic AI in use at 23% of firms currently, projected 74% within two years; only approximately 20% have mature governance models for autonomous agents. https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html Palo Alto Networks, Unit 42 incident response reporting and 2026 AI security forecasts (2025–2026). Formally classifies autonomous AI agents as a new insider threat category; documents incidents of agents retaining access after human role changes; launched Prisma AIRS 3.0 for agent discovery and runtime protection. https://www.paloaltonetworks.com/company/press/2025/palo-alto-networks-forecasts-6-predictions-on-securing-the-new-ai-economy-for-2026 Orchid Security, Extending IAM for Agent AI (May 2026). Introduces delegated identity architecture requiring explicit delegation chains and treating agent memory as an auditable privileged action. https://www.orchid.security/reports/how-to-extend-iam-for-agent-ai Cloud Security Alliance, Non-human identity and agentic AI governance whitepaper (May 2026). Calls for dynamic agent inventories integrated with identity management, lifecycle governance including memory disposition on decommissioning, and real-time registries for spawned sub-agents. https://labs.cloudsecurityalliance.org/research/csa-whitepaper-nonhuman-identity-agentic-ai-governance-v1-cs/ ServiceNow, Vancouver release notes (November 2023). AI agents retain workflow context across IT, HR, and security modules; three healthcare system deployments include automatic privilege expiration after 30 days of inactivity. Anthropic, Computer use API documentation (October 2024). Long-running agents with file system and browser tool access; broad scope that static authority matrices cannot track. Technical readers can find detailed customer metrics and benchmarks in the original announcements linked above.

bottom of page