Building Agents with Model Context Protocol
Shopkick partners with Christopher Daniel to achieve MCP (Model Context Protocol)
About Shopkick
Shopkick is a mobile app that rewards people for the everyday things they already do while shopping. Users can earn points, called “kicks,” by walking into participating stores, scanning products, browsing offers, or making purchases. Those points can then be exchanged for gift cards, giving shoppers a small reward for discovering new products and engaging with brands.
The platform was originally launched in the U.S. and is now part of Trax, a retail technology company founded in Singapore. Shopkick was built around a simple idea: make shopping a little more interactive for consumers while helping brands reach people at the moment they’re deciding what to buy.
For retailers and product brands, Shopkick provides a way to drive store visits, increase product discovery, and encourage trial without relying only on traditional advertising or discounts. By connecting digital engagement with in-store behavior, the platform gives brands a better understanding of how shoppers actually move through the buying process.
Shopkick’s North Star Goal
The goal of the platform is to turn fragmented operational data into immediate clarity during outages. By combining code changes, infrastructure configurations, and live performance metrics into a single AI-driven workflow, the system gives teams a faster path from detection to resolution. Instead of relying on manual investigation across multiple dashboards, engineers receive a consolidated diagnosis, rollback strategy, and verified supporting data in one place.
Over time, the North Star is to reduce the average Mean Time to Resolution (MTTR) across critical incidents by automating the investigative process and giving leadership a clear understanding of what happened, why it happened, and how it should be fixed.
Evolution of Client's stack with Christopher Daniel
The transformation of the client's technology stack led by Christopher Daniel represented a paradigm shift from a limited retrieval system to an intelligent orchestration engine. Previously, ShopBack Singapore had utilized a basic Retrieval Augmented Generation application built on top of siloed internal databases. The front-end interface could only answer basic statistical questions based on available data, completely lacking the intelligence required to combine proprietary internal data with live external APIs and unstructured sources. This static approach prevented the generation of real-time diagnostic outputs and required decision making to depend heavily on manual analysis and static reporting.
By implementing a multi agent framework, Christopher Daniel transformed the system into a continuously evolving, context-aware decision support engine. Once deployed, it dynamically ingested live system metrics, code commits, cloud configurations, and network routing events. With this comprehensive data ingestion, the platform provided stakeholders with a massive variety of strategic assets. It successfully generated complete incident summary reports, developed detailed rollback strategies, and suggested code fixes perfectly aligned with current system errors. Ultimately, it outputted a finalized executive dashboard that told the internal leadership team exactly why an outage occurred and the optimal steps to resolve it.
Architecture
The revised architecture we built was entirely driven by a LangGraph Supervisor Node acting as the central brain to trigger workflows and manage multiple specialized agents. We utilized the powerful Claude 3.5 Sonnet model to drive the core reasoning, synthesis, and iterative prompting capabilities required by the client. A key architectural decision was to build heavily on standardized Model Context Protocol tools to ensure seamless data pipelines.
We divided the architecture into five core operational areas:
- The Data Origins: The physical reality of software engineers deploying code, global traffic hitting the network, and application servers generating operational metrics.
- Live Enterprise Tools: The recording layer where GitHub tracks code changes, Datadog APM logs CPU spikes, AWS EC2 hosts the active cloud, and Cloudflare Edge monitors web traffic.
- Standardized Connectors: This is the critical Model Context Protocol layer. Instead of custom APIs, we deployed a GitHub MCP Server, Datadog MCP Server, Terraform MCP Server, and Cloudflare MCP Server. These act as universal translators for the AI.
- Multi Agent Orchestration: Hosted on a FastAPI and LangServe backend, a LangGraph Supervisor Node acts as the Chief Incident Agent. It delegates tasks to specialized LangChain ReAct Agents. A Code Specialist Agent accesses the GitHub and Datadog MCPs, while an Infrastructure Specialist Agent queries the Terraform and Cloudflare MCPs.
- The Executive Dashboard: A React Web Application that allows the executive user in Singapore to ask plain English questions. The dashboard displays the AI's final root cause analysis alongside clickable, verified data citations directly from the MCP servers.
Implementation
Because the client operations were centered in Singapore, all AWS infrastructure resources were hosted in the ap-southeast-1 region to ensure low latency. We leveraged AWS EC2 instances, Lambda, and complex step functions to ensure scalability and secure integration with the client's existing virtual private clouds. The MCP servers were individually Dockerized and orchestrated using Kubernetes.
A primary challenge was managing the rate limits of external third-party feeds and APIs. We implemented asynchronous queuing and exponential backoff strategies to prevent throttling. Additionally, keeping the latency low while the backend consolidated insights from the Datadog MCP and live streaming performance required aggressive query optimization and caching mechanisms.
This best-in-class implementation required a cross functional team of 9 professionals. The team consisted of 2 Data Engineers to build pipelines, 3 AI/ML Engineers focusing on LangGraph orchestration and MCP integrations, 1 AWS Solutions Architect, 1 React Frontend Developer, 1 Quality Assurance Tester, and 1 Technical Project Manager.
Evaluation of the Model
To ensure the model functioned as a reliable enterprise tool rather than a novelty, we established rigorous evaluation protocols using industry standard frameworks for Large Language Model applications. We utilized the RAGAS framework to quantitatively measure Context Relevance, Groundedness, and Answer Relevance.
To measure Groundedness, we employed an LLM as a Judge pipeline. A secondary evaluation model automatically audited the Chief Agent's finalized incident reports to detect hallucination rates, ensuring every server status or code citation could be traced directly back to the Datadog or GitHub MCP servers. Operational performance was continuously tracked using Mean Time to Resolution metrics to measure the precision of the AI's diagnostics. Because the solution supports iterative prompting and constraint-based reasoning, we heavily tracked the approval rate at the Human in the Loop checkpoint. By monitoring how often the human approver in the React dashboard had to request an AI revision versus granting immediate approval for the diagnostics package, we established a clear baseline for the model's practical utility.
Future Priorities
The current implementation is just a minor piece in a major puzzle. While the multi agent system is fully capable of working independently, it is fundamentally designed to integrate with larger plans that are going to be implemented in the future. To support this expanded vision and ensure the system can scale reliably, the technical plan must prioritize the following foundational pillars:
- Robust API Management and Scaling: As more enterprise applications leverage the REST API feature, establishing an enterprise grade API Gateway will be mandatory. This will manage heavy traffic, enforce strict authentication protocols, and provide necessary load balancing to prevent the core system from being bottlenecked by requests from other internal tools.
- Modular Agent Expansion: Future iterations will likely require new specialized components, such as a Security Compliance Agent or a Billing Agent. The architecture must maintain strict decoupling, so the LangGraph Supervisor can dynamically register new agents without disrupting foundational data fetchers via MCP servers.
- Advanced LLMOps Pipelines: Implementing continuous integration and deployment for prompt templates will be crucial. Establishing an LLMOps framework will allow the engineering team to version control the AI's reasoning logic, ensuring that updates to the underlying Claude model do not degrade the contextual accuracy of the system.
- Zero Trust Data Governance: Opening the orchestration engine to other internal tools via the REST API necessitates rigorous, granular access controls. We must ensure that downstream applications only receive the data they are authorized to view, safeguard sensitive metrics handled by components like the internal agents and ensure strict compliance with enterprise security policies.