We spend a lot of time worrying about AI hallucinations. The model confidently says something that is not true. We add guardrails, evaluation loops, grounding pipelines. We are very serious about this problem.
Meanwhile, we are doing the exact same thing with AI concepts.
Every few weeks, a new idea arrives. We confidently pivot to it. We abandon the last one. We repeat. If an LLM did this, we would call it hallucinating. When we do it, we call it staying current.
Speed of AI adoption is not the competitive advantage. Quality of understanding is.
A couple of words about me:
I am Sunny, Lead Product Manager for Cloud Integration at SAP Integration Suite, based at SAP Labs India. I have spent 16+ years at SAP, starting as an intern to developer in Java, JavaScript, UI5, etc. before moving into product management.
The views expressed in this post are my own, based on observations and patterns I have seen in this space. They do not represent the official position of SAP.
The AI Concept Treadmill:
Figure 1 – Running hard on every new AI concept. Never quite arriving.
It started, as most things do, with a paper. In 2020, researchers at Facebook AI Research formalized Retrieval-Augmented Generation (RAG). The idea was elegant: instead of baking all knowledge into a model at training time, retrieve relevant context at inference time, then generate. Enterprises picked this up seriously around 2023, when the ChatGPT adoption wave made connecting AI to internal data a boardroom priority overnight.
Then came the fine-tuning wave of mid-2023. Why retrieve external knowledge when you could train the model on your data directly? Enterprises started debating: fine-tune or RAG? It was a legitimate question and many teams were still working through it when the next wave hit.
Then came a one-two punch on context windows. In November 2023, OpenAI’s GPT-4 Turbo arrived with a 128k token context window. Three months later, in February 2024, Gemini 1.5 raised that to one million tokens. Together, they triggered a genuine question: if you can fit an entire corpus into a single prompt, do you still need a retrieval pipeline? Teams that had invested months in RAG found themselves revisiting their assumptions which, honestly, is exactly the right thing to do when the landscape shifts.
Simultaneously, the autonomous agent thread was already running. AutoGPT launched in March 2023 and became the top trending GitHub repository within days and the collapse was nearly as fast when it could not reliably complete real-world tasks. In 2024, improved tool-calling revived the conversation, this time with more structure.
Late 2024 brought reasoning models: OpenAI’s o1, then DeepSeek R1 in early 2025. A different kind of intelligence — slower, more deliberate, trained to think through steps before answering.
Then Anthropic released the Model Context Protocol (MCP) in November 2024. A standard for how AI agents call tools and data sources. Within weeks: Isn’t this just JSON-RPC? Just REST APIs with extra steps? The debate was real, heated, and productively confused.
Then, in April 2025, Google introduced Agent-to-Agent protocol (A2A). Understandably, the community asked: how does this relate to MCP? Do we need both? Which one should we build on? These are exactly the right questions, the problem is when we try to answer them before we have read either spec carefully.
And now, in 2025, the enterprise conversation is settling slowly, carefully on agentic workflows: structured, observable, policy-enforced multi-step AI, with human oversight where it matters. Governed autonomy.
Sound familiar? That feeling of perpetual catch-up — of never quite finishing one thing before the next arrives — that is the treadmill. And the problem is not the pace of AI innovation. The problem is our reaction to it.
Eight distinct concepts across five years. The last four arrived in eighteen months. The AI Concept Tradmill is not slowing down.
Where the Real Confusion Lives:
Figure 2 – Four pairs that confuses us.
1. RAG vs. Long Context Windows
The ‘RAG is dead’ conclusion spread quickly and understandably so. A one-million-token context window sounds like it renders retrieval pipelines redundant. But RAG and long context windows solve overlapping problems in fundamentally different ways. The context window is static at inference time: you load it once, it does not update mid-conversation. RAG is dynamic, it retrieves only what is relevant, at query time, from a corpus that can be updated continuously. The context window is also expensive and slow to process at scale. For an enterprise with millions of product records and transaction histories, retrieval is not optional — it is the only practical path. RAG is not dead. It is complementary.
A well-designed enterprise AI system often uses long context for deep reasoning over a curated set of documents, and RAG for real-time retrieval from a live, updating corpus. The question is not which one, it is understanding when to use which, and at what cost.
2. MCP vs. APIs
The ‘isn’t this just APIs?’ reaction is understandable, but it conflates two fundamentally different things. A REST or OData API is a static contract: you define endpoints, methods, and payloads, and a caller that already knows the interface invokes it. MCP is a dynamic capability protocol designed specifically for AI agents: a model can discover what tools exist, understand what each one does from its description, decide which one to call based on the task at hand, invoke it, and handle the result, all at runtime, without any hardcoded knowledge of the interface. The difference is not syntax. It is intent. APIs were designed for developers integrating systems. MCP was designed for AI agents acting autonomously.
In fact, APIs are often what fuel MCP tools under the hood. An MCP Server wraps existing API capabilities and exposes them in a way that AI agents can discover and invoke natively. It is not APIs vs MCP. It is APIs plus MCP.
Did You Know: SAP Integration Suite’s MCP Gateway which acts as a governed entry point across MCP Servers that provides governance, observability, and policy enforcement built in, making AI agents truly enterprise-ready.
3. Autonomous Agents vs. Agentic Workflows
AutoGPT’s rise and fall is the clearest case study we have. The promise was full autonomy: set the goal, let the agent run. Real-world tasks require judgment calls and error recovery that fully autonomous agents in 2023 simply could not handle reliably. What enterprises are deploying successfully in 2025 is agentic workflows: structured, multi-step AI processes with defined scope, human-in-the-loop checkpoints, and full observability.
Agentic workflows are how enterprises deploy agent capabilities today. Fully autonomous agents are where the space is heading. Understanding the difference tells you what to build now versus what to watch and that is a much more useful frame than ‘autonomous agents don’t work.’
4. A2A vs. MCP
These two protocols operate at different layers of the same architecture, not in competition with each other. MCP governs how an AI agent calls a tool or data source, the agent-to-tool layer. A2A governs how AI agents communicate with each other, the agent-to-agent layer. In a mature multi-agent enterprise architecture, you will likely need both.
Think of it this way: A2A handles how agents coordinate, and MCP handles what each agent can do. One governs the conversation between agents, the other governs the capability each agent brings to it. The question is not which one wins, it is understanding what problem each one solves so you can use them together.
What This Pattern Costs Enterprises
This is not abstract. The reactive adoption pattern has real business consequences, and I have seen them described consistently across enterprise AI discussions throughout 2024 and 2025.
I have seen this pattern play out consistently across enterprise AI discussions throughout 2024 and 2025 and I include myself in it. Teams that invested in RAG pipelines, revisited them when long-context windows arrived, rebuilt with a different approach, and then realized their original use case actually needed retrieval dynamics — those teams lost months and budget. More significantly: they lost momentum and stakeholder confidence at exactly the point when they needed both.
A common pattern in 2024 was AI initiatives stuck in pilot purgatory — proofs of concept that could not graduate to production because the underlying approach kept shifting before it stabilized.
There is also the subtler cost: opportunity cost. While the organization is chasing the new concept, the previous one, which was actually the right fit, sits unimplemented. The enterprise that patiently built a well-governed RAG pipeline in 2023 and refined it through 2024 is often in a stronger position today than the one that pivoted three times chasing the frontier.
Chasing every new wave is itself a kind of hallucination.
The Grounded Adoption Framework
Figure 3 – Five steps from reactive adoption to grounded decision-making.
- Understand. Go back to the source — the original spec, paper, or announcement. Not because summaries are bad, but because the design intent rarely survives the translation to a headline. The MCP specification is publicly available and readable in under an hour. That hour is worth it.
- Correlate. Map the new concept against what you already know. Is this net new? An evolution? A specialization? A renaming? Most new AI concepts are roughly 70% familiar and 30% genuinely new. Your job is to find that 30%, that is where the real learning is.
- Unlearn deliberately. Identify what prior assumption this new concept actually invalidates, if any. Not every new concept requires discarding something old. ‘MCP provides a standardized invocation layer that replaces ad-hoc direct API calls for agent tool use’ is a legitimate unlearn. ‘RAG is dead’ is a false unlearn. Do not do it.
- Relearn properly. Update your mental model with the new concept correctly placed. Resist the temptation to layer it on top of the old one without resolving the overlap, that is how well-intentioned explanations like ‘MCP is basically just APIs’ take root, and once they do they are hard to correct in a room full of stakeholders.
- Decide consciously. Now and only now, evaluate whether this concept is relevant to your enterprise context, your maturity level, your specific use case. Understanding a concept deeply and concluding it is not right for you today is a completely valid, intelligent, professional outcome.
Important Points to Consider
- Not every new concept requires a decision. Understanding is not the same as adopting. You can understand A2A deeply and conclude it is not relevant to your enterprise today. That is a valid, informed, professional outcome, not a missed opportunity.
- The AI Concept Treadmill will not slow down. The right response is not to wait for AI to stabilize, it will not. The right response is to build the muscle of grounded evaluation so you can process new concepts faster and more accurately over time.
- Your integration layer is your grounding layer. For enterprise AI, governed, real-time access to business context via secured APIs, event streams, and integration middleware with proper security, monitoring, audit, and policy enforcement is what separates reliable enterprise AI from hallucinating enterprise AI. Ground your systems. Ground yourself.
Conclusion
The enterprises that will lead with AI are not the ones who adopted the most concepts. They are the ones who understood the right ones well enough to implement them well.
The LLMs are getting better at not hallucinating. The question is: are we?
Now let’s hear from you:
I’d love to hear your thoughts on navigating the AI adoption landscape. Share your insights by answering these questions in the comments below:
- Which AI concept caused the most confusion or reactive pivoting in your organization and how did you course-correct?
- What is your personal process for evaluating a new AI concept before deciding to adopt it for your enterprise?
Let’s start a conversation and learn from each other’s perspectives!
We spend a lot of time worrying about AI hallucinations. The model confidently says something that is not true. We add guardrails, evaluation loops, grounding pipelines. We are very serious about this problem.Meanwhile, we are doing the exact same thing with AI concepts.Every few weeks, a new idea arrives. We confidently pivot to it. We abandon the last one. We repeat. If an LLM did this, we would call it hallucinating. When we do it, we call it staying current.Speed of AI adoption is not the competitive advantage. Quality of understanding is.A couple of words about me:I am Sunny, Lead Product Manager for Cloud Integration at SAP Integration Suite, based at SAP Labs India. I have spent 16+ years at SAP, starting as an intern to developer in Java, JavaScript, UI5, etc. before moving into product management. The views expressed in this post are my own, based on observations and patterns I have seen in this space. They do not represent the official position of SAP.The AI Concept Treadmill: Figure 1 – Running hard on every new AI concept. Never quite arriving.It started, as most things do, with a paper. In 2020, researchers at Facebook AI Research formalized Retrieval-Augmented Generation (RAG). The idea was elegant: instead of baking all knowledge into a model at training time, retrieve relevant context at inference time, then generate. Enterprises picked this up seriously around 2023, when the ChatGPT adoption wave made connecting AI to internal data a boardroom priority overnight.Then came the fine-tuning wave of mid-2023. Why retrieve external knowledge when you could train the model on your data directly? Enterprises started debating: fine-tune or RAG? It was a legitimate question and many teams were still working through it when the next wave hit.Then came a one-two punch on context windows. In November 2023, OpenAI’s GPT-4 Turbo arrived with a 128k token context window. Three months later, in February 2024, Gemini 1.5 raised that to one million tokens. Together, they triggered a genuine question: if you can fit an entire corpus into a single prompt, do you still need a retrieval pipeline? Teams that had invested months in RAG found themselves revisiting their assumptions which, honestly, is exactly the right thing to do when the landscape shifts.Simultaneously, the autonomous agent thread was already running. AutoGPT launched in March 2023 and became the top trending GitHub repository within days and the collapse was nearly as fast when it could not reliably complete real-world tasks. In 2024, improved tool-calling revived the conversation, this time with more structure.Late 2024 brought reasoning models: OpenAI’s o1, then DeepSeek R1 in early 2025. A different kind of intelligence — slower, more deliberate, trained to think through steps before answering.Then Anthropic released the Model Context Protocol (MCP) in November 2024. A standard for how AI agents call tools and data sources. Within weeks: Isn’t this just JSON-RPC? Just REST APIs with extra steps? The debate was real, heated, and productively confused.Then, in April 2025, Google introduced Agent-to-Agent protocol (A2A). Understandably, the community asked: how does this relate to MCP? Do we need both? Which one should we build on? These are exactly the right questions, the problem is when we try to answer them before we have read either spec carefully.And now, in 2025, the enterprise conversation is settling slowly, carefully on agentic workflows: structured, observable, policy-enforced multi-step AI, with human oversight where it matters. Governed autonomy.Sound familiar? That feeling of perpetual catch-up — of never quite finishing one thing before the next arrives — that is the treadmill. And the problem is not the pace of AI innovation. The problem is our reaction to it.Eight distinct concepts across five years. The last four arrived in eighteen months. The AI Concept Tradmill is not slowing down.Where the Real Confusion Lives: Figure 2 – Four pairs that confuses us.1. RAG vs. Long Context WindowsThe ‘RAG is dead’ conclusion spread quickly and understandably so. A one-million-token context window sounds like it renders retrieval pipelines redundant. But RAG and long context windows solve overlapping problems in fundamentally different ways. The context window is static at inference time: you load it once, it does not update mid-conversation. RAG is dynamic, it retrieves only what is relevant, at query time, from a corpus that can be updated continuously. The context window is also expensive and slow to process at scale. For an enterprise with millions of product records and transaction histories, retrieval is not optional — it is the only practical path. RAG is not dead. It is complementary.A well-designed enterprise AI system often uses long context for deep reasoning over a curated set of documents, and RAG for real-time retrieval from a live, updating corpus. The question is not which one, it is understanding when to use which, and at what cost.2. MCP vs. APIsThe ‘isn’t this just APIs?’ reaction is understandable, but it conflates two fundamentally different things. A REST or OData API is a static contract: you define endpoints, methods, and payloads, and a caller that already knows the interface invokes it. MCP is a dynamic capability protocol designed specifically for AI agents: a model can discover what tools exist, understand what each one does from its description, decide which one to call based on the task at hand, invoke it, and handle the result, all at runtime, without any hardcoded knowledge of the interface. The difference is not syntax. It is intent. APIs were designed for developers integrating systems. MCP was designed for AI agents acting autonomously.In fact, APIs are often what fuel MCP tools under the hood. An MCP Server wraps existing API capabilities and exposes them in a way that AI agents can discover and invoke natively. It is not APIs vs MCP. It is APIs plus MCP. Did You Know: SAP Integration Suite’s MCP Gateway which acts as a governed entry point across MCP Servers that provides governance, observability, and policy enforcement built in, making AI agents truly enterprise-ready.3. Autonomous Agents vs. Agentic WorkflowsAutoGPT’s rise and fall is the clearest case study we have. The promise was full autonomy: set the goal, let the agent run. Real-world tasks require judgment calls and error recovery that fully autonomous agents in 2023 simply could not handle reliably. What enterprises are deploying successfully in 2025 is agentic workflows: structured, multi-step AI processes with defined scope, human-in-the-loop checkpoints, and full observability. Agentic workflows are how enterprises deploy agent capabilities today. Fully autonomous agents are where the space is heading. Understanding the difference tells you what to build now versus what to watch and that is a much more useful frame than ‘autonomous agents don’t work.’4. A2A vs. MCPThese two protocols operate at different layers of the same architecture, not in competition with each other. MCP governs how an AI agent calls a tool or data source, the agent-to-tool layer. A2A governs how AI agents communicate with each other, the agent-to-agent layer. In a mature multi-agent enterprise architecture, you will likely need both. Think of it this way: A2A handles how agents coordinate, and MCP handles what each agent can do. One governs the conversation between agents, the other governs the capability each agent brings to it. The question is not which one wins, it is understanding what problem each one solves so you can use them together.What This Pattern Costs EnterprisesThis is not abstract. The reactive adoption pattern has real business consequences, and I have seen them described consistently across enterprise AI discussions throughout 2024 and 2025.I have seen this pattern play out consistently across enterprise AI discussions throughout 2024 and 2025 and I include myself in it. Teams that invested in RAG pipelines, revisited them when long-context windows arrived, rebuilt with a different approach, and then realized their original use case actually needed retrieval dynamics — those teams lost months and budget. More significantly: they lost momentum and stakeholder confidence at exactly the point when they needed both.A common pattern in 2024 was AI initiatives stuck in pilot purgatory — proofs of concept that could not graduate to production because the underlying approach kept shifting before it stabilized.There is also the subtler cost: opportunity cost. While the organization is chasing the new concept, the previous one, which was actually the right fit, sits unimplemented. The enterprise that patiently built a well-governed RAG pipeline in 2023 and refined it through 2024 is often in a stronger position today than the one that pivoted three times chasing the frontier.Chasing every new wave is itself a kind of hallucination.The Grounded Adoption Framework Figure 3 – Five steps from reactive adoption to grounded decision-making. Understand. Go back to the source — the original spec, paper, or announcement. Not because summaries are bad, but because the design intent rarely survives the translation to a headline. The MCP specification is publicly available and readable in under an hour. That hour is worth it. Correlate. Map the new concept against what you already know. Is this net new? An evolution? A specialization? A renaming? Most new AI concepts are roughly 70% familiar and 30% genuinely new. Your job is to find that 30%, that is where the real learning is. Unlearn deliberately. Identify what prior assumption this new concept actually invalidates, if any. Not every new concept requires discarding something old. ‘MCP provides a standardized invocation layer that replaces ad-hoc direct API calls for agent tool use’ is a legitimate unlearn. ‘RAG is dead’ is a false unlearn. Do not do it. Relearn properly. Update your mental model with the new concept correctly placed. Resist the temptation to layer it on top of the old one without resolving the overlap, that is how well-intentioned explanations like ‘MCP is basically just APIs’ take root, and once they do they are hard to correct in a room full of stakeholders. Decide consciously. Now and only now, evaluate whether this concept is relevant to your enterprise context, your maturity level, your specific use case. Understanding a concept deeply and concluding it is not right for you today is a completely valid, intelligent, professional outcome.Important Points to Consider Not every new concept requires a decision. Understanding is not the same as adopting. You can understand A2A deeply and conclude it is not relevant to your enterprise today. That is a valid, informed, professional outcome, not a missed opportunity. The AI Concept Treadmill will not slow down. The right response is not to wait for AI to stabilize, it will not. The right response is to build the muscle of grounded evaluation so you can process new concepts faster and more accurately over time. Your integration layer is your grounding layer. For enterprise AI, governed, real-time access to business context via secured APIs, event streams, and integration middleware with proper security, monitoring, audit, and policy enforcement is what separates reliable enterprise AI from hallucinating enterprise AI. Ground your systems. Ground yourself.ConclusionThe enterprises that will lead with AI are not the ones who adopted the most concepts. They are the ones who understood the right ones well enough to implement them well.The LLMs are getting better at not hallucinating. The question is: are we? Now let’s hear from you:I’d love to hear your thoughts on navigating the AI adoption landscape. Share your insights by answering these questions in the comments below:Which AI concept caused the most confusion or reactive pivoting in your organization and how did you course-correct?What is your personal process for evaluating a new AI concept before deciding to adopt it for your enterprise?Let’s start a conversation and learn from each other’s perspectives! Read More Technology Blog Posts by SAP articles
#SAPCHANNEL