MCP + Skill in Practice: A Guide to Building an Automated Enterprise Regulatory Risk Alert System

Build an end-to-end automated regulatory risk alert system using MCP and Agent Skill architecture.
This guide demonstrates how to combine MCP (Model Context Protocol) with Agent Skill to build a complete enterprise regulatory risk monitoring system. It covers structured enterprise profiling for applicability judgment, six evidence thresholds to ensure reliability, a three-layer report structure (risk description, applicability argument, actionable measures), and automated delivery via Feishu or email — turning raw regulatory data into trustworthy, actionable intelligence.
In today's increasingly complex regulatory landscape, enterprises face a daily deluge of government announcements, administrative penalty cases, and local regulation updates. Being able to see a policy doesn't mean knowing what to do about it. Based on a real-world Agent Skill case shared by the Bazhuayun team, this article breaks down how to use MCP (Model Context Protocol) to create a complete closed loop from "data collection → professional judgment → evidence trail → delivery and notification," turning an AI Agent into a true regulatory risk sentinel for enterprises.
MCP and Agent Skill: Understanding Two Core Concepts
Before diving into the case study, it's important to clarify two key concepts.
MCP (Model Context Protocol) is a protocol standard open-sourced by Anthropic in late 2024, designed to provide large language models with a unified way to access external tools and data sources. Before MCP, every AI Agent that needed to call external services (such as database queries, web scraping, or API calls) required a custom adapter layer, resulting in high development costs and poor reusability. MCP's core idea is to define a standardized "server-client" communication protocol that allows Agents to discover, invoke, and manage various external capabilities through a unified interface. In this case study, MCP serves as the bridge between the AI Agent and the Bazhuayu scraping platform — the Agent doesn't need to understand the specific implementation details of the crawler; it simply starts tasks, queries status, and exports data through the MCP interface, achieving standardized encapsulation and reuse of data acquisition capabilities.
Agent Skill refers to packaging a complete business capability as a standardized module that can be invoked by an AI Agent. Unlike simple prompt templates, a Skill encompasses constraints across four dimensions: input specifications (what parameters are needed), processing logic (what steps to follow), quality standards (what conditions the output must meet), and delivery methods (how results reach the user). This concept borrows from the "microservices" philosophy in software engineering — each Skill is an independent, composable capability unit, and different Agents can call different Skill combinations depending on the scenario. In this case, regulatory risk analysis is packaged as a Skill, meaning that regardless of which Agent runs it or on what platform, as long as the correct parameters are provided, a risk report of consistent quality will be produced.
Why Regulatory Information Is Hard to Turn Into Enterprise Decisions
Enterprises inevitably face compliance risks in their operations, and compliance and legal teams need to continuously monitor penalty lists and policy announcements across various regulatory agencies and government information platforms. The problem is that this information is vast in volume and frequently updated, making manual curation extremely difficult.
The real pain points mentioned in the case can be summarized across three levels:
- Seeing a policy ≠ knowing what to do. More regulatory information doesn't automatically improve risk management quality. The real difficulty lies in determining whether a document applies to a specific enterprise — whether the region is covered, whether the entity type matches, whether licenses and products are affected, and whether the document has taken effect.
- Risk extraction is hard. Enterprises don't care how long a policy document is; they care about what prohibitions, reporting obligations, rectification requirements, or consumer protection duties it triggers, and what consequences arise from inaction.
- Implementation is hard. Reports without responsible parties, deadlines, deliverables, and review criteria easily become one-time reading material.

These three challenges are exactly what this Skill aims to solve. It doesn't just answer "what happened" but continues to answer three critical questions: What is the risk? Why is it relevant to this enterprise? What should be done next?
Enterprise Profile: Making Applicability Judgments Explicit
In traditional compliance work, lawyers or compliance officers mentally compare enterprise characteristics against regulatory applicability conditions when reading regulatory documents — a highly experience-dependent implicit judgment process. The key innovation in this case is introducing a structured "Enterprise Profile" that makes this implicit judgment explicit and parameterized.
An enterprise profile typically includes dimensions such as operating regions, license types held, main products and services, customer segments (individual/institutional), sales channels (online/offline), data processing scope, and technology outsourcing arrangements. When these characteristics are passed into the system in a structured format, the AI can compare them item by item against the scope of regulatory documents, transforming the core question "Is this regulation relevant to me?" into a computable logical judgment rather than relying on the model's "intuition."
The Three-Layer Structure of a High-Quality Regulatory Risk Report
In the demonstration, 108 items were collected from two regulatory websites. The final report flagged 3 high-risk items and 2 items for observation, each accompanied by assessment rationale and original text excerpts. The report's core value lies in its three-layer progressive structure:
Layer 1: Precise Risk Description
A risk point is not a document title, nor a lengthy policy summary. It's a single sentence that clearly states what obligation, regulatory concern, or conditional risk the enterprise has triggered. For example, "The online sales process for individual customers lacks the suitability confirmation checkpoint required by the new regulation" — rather than the vague "consumer protection risk in sales process."
Layer 2: Applicability Argument — Why It's Relevant to This Enterprise
This layer connects the matched enterprise profile fields (operating region, license, product, customer type, sales channel, data processing, technology outsourcing, etc.) with official sources, publication dates, original text excerpts, and applicability logic, so readers can understand why a risk is relevant without having to re-read the entire document.
Layer 3: Actionable Implementation Measures
Recommendations cannot stop at empty phrases like "continue monitoring" or "strengthen management." Each significant risk should, where possible, include a responsible party, deadline, deliverable, and review evidence. For example: "Compliance to lead a product suitability review within 10 business days, delivering a gap analysis checklist, confirmed by the business owner."
One noteworthy detail: the report displays risk level and judgment confidence separately, refusing to classify something as high risk simply because the model uses strong language. This separation stems from the classic risk management framework — Risk = Impact × Likelihood — but takes on special significance in the context of AI-generated content. Large language models naturally tend toward definitive-sounding language, even when their actual evidence is insufficient. Without an independent confidence dimension, the model's "confident tone" can easily be misread as "a definitive high-risk judgment." By explicitly requiring the system to lower confidence markings when evidence is insufficient — even if the risk level itself might be high — items get flagged as "pending verification" rather than going straight into alerts, effectively countering the erosion of professional judgment by AI hallucination.
The Five-Role Separation of Responsibilities Architecture
To make judgments trustworthy, clear boundaries are essential. The case splits the entire system into five roles, each with distinct responsibilities:
- LLM (Large Language Model): Responsible for semantic understanding, processing Chinese text with inconsistent field structures. But it cannot replace official sources, nor score freely based on linguistic intuition. The Skill places the LLM inside a "rule frame" — only after passing source, validity, and applicability conditions is the model allowed to generate risk explanations.
- Agent: Responsible for process — asking for necessary inputs, launching tasks, polling status. Any missing critical input defaults to blocking continuation; collection failures or unreadable data return a clear blocking status rather than forcing out a report.
- MCP + Collection Tasks: Responsible for data. Users don't need the Agent to write scrapers on the fly; they simply select pre-packaged collection tasks or templates, and MCP handles launching, querying, and exporting. Particularly suitable for websites with irregularly updated announcements, regulations, and administrative penalties.
- Skill: Responsible for standards. It codifies four types of standards — input standards, evidence standards, output standards, and publication standards — so different Agents can call it without reinventing the process, and professional boundaries won't be lost due to prompt variations.
- Enterprise Personnel: Still retain the right to confirm and review. The report is decision support, not legal advice.

This separation of responsibilities is critical: the system won't automatically determine that an administrative penalty means the enterprise violated the law, won't pretend applicability is certain when the enterprise profile is incomplete, and won't generate high-risk conclusions when evidence thresholds haven't been met.
Core Mechanism: Six Evidence Thresholds Explained
Once data credibility is established, the system doesn't immediately assign scores. Instead, it first determines whether the policy matches the enterprise. This matching isn't keyword-based — a document mentioning "financial institutions" doesn't mean all related enterprises are covered. It also depends on entity definitions, licenses, business activities, and effective dates.
A special note on the use of administrative penalty information: penalty decisions published by regulatory agencies typically contain findings of illegal conduct, applicable legal provisions, and penalty types and amounts. For enterprises, these cases serve three main purposes: first, they reveal the regulator's current enforcement priorities and focus areas; second, they demonstrate the actual consequences of specific violations (fine amounts, license revocations, etc.), helping enterprises assess risk exposure; third, through the factual descriptions in cases, enterprises can reverse-engineer whether they have similar business practices. However, as emphasized in this case design, administrative penalty cases can only be used to identify regulatory concerns and potential consequences — they cannot be used to infer that the enterprise has already violated the law, because each penalty has its specific factual context and investigation process.
Only after applicability is established does the system proceed to risk extraction and measure formulation. The reliability of all this rests on six evidence thresholds:
- Source Authenticity: Is it from an official agency? Can the link be opened? Is the full text complete?
- Validity Status: Currently effective, future effective, expired, open for comments, or general notice?
- Applicability: Does it match standard parameters like region and entity type?
- Nature of Obligation: Prohibition, mandatory requirement, rectification, or internal control?
- Risk Consequences: Licensed operations, administrative penalties, etc.?
- Enterprise Exposure: Does the enterprise profile actually contain business activities that trigger this obligation?

Only when official source, complete text, clear applicability basis, and matched enterprise profile fields are all present simultaneously can information become a source for severe or high-risk classification. Additionally, "confidence" operates as an independent axis — when key text, entity definitions, or enterprise facts are insufficient, even a match cannot enter high-risk status, preventing the model from packaging uncertainty as high risk.
End-to-End Operations: The Eight-Step Closed Loop
The demonstration used both Warp (work body) and Codex pipelines for operations. The key to the entire workflow is that users only need to provide four core parameters:
- Bazhuayu MCP API Key: Tells the Skill which account's tasks to collect
- Enterprise Profile: Lets the Skill understand what the enterprise looks like to judge applicability
- Collection Task: Specifies the data source
- Delivery Channel: Feishu or email
Regardless of how the Agent probes for information, users only need to provide details around these four points. The system guides users through configuring Feishu CLI or email SMTP authorization codes.
Regarding the technical implementation of delivery channels: Feishu push notifications are implemented through the Feishu Open Platform's bot API, requiring the target user's openid (the user's unique identifier on the Feishu platform), which can be obtained by querying through the user's phone number or email in the Feishu Open Platform's API debugging console. Email push is based on the SMTP protocol, where users need to provide their email's authorization code rather than their login password — this is because mainstream email service providers (such as QQ Mail, NetEase Mail) require third-party applications to use specially generated app-specific passwords instead of login passwords for SMTP authentication for security purposes. This design allows the system to achieve automated push notifications without requiring users to expose their primary email passwords.
After task parsing, the system records the run ID, initiates parallel execution, polls, exports full data, and generates a snapshot for the current run. The report only uses the current snapshot — each time the task ID or task name changes, a completely new report is generated, avoiding redundancy from historical data duplication. For collection results with inconsistent field names (some called "full text," others called "content"), the Skill dynamically identifies and preserves original fields, ultimately categorizing each piece of information into one of four outcomes: enters risk, enters observation, enters pending verification, or determined not applicable.

The final report is generated as a single-file HTML and pushed via Feishu or email. In the demonstration, the system successfully identified 2 severe risks, several high risks, and observation items, sending risk points, reasoning, and implementation measures together. This Skill can also be packaged as a scheduled task — for example, automatically executing collection every morning and pushing to a designated email.
User Value and Extensibility
Compared to the traditional process (collecting links → manual screening, reading, and summarizing → producing lengthy interpretations with incomplete execution accountability and evidence trails), the Skill approach first performs applicability judgment, then extracts risk points and measures, placing evidence fingerprints and delivery receipts within the same closed loop. Its value can be summarized in three points:
- Faster: Shortens the path from data to judgment
- More Accurate: Filters based on enterprise profile rather than generic interpretation
- More Actionable: Keeps risk, rationale, and action in sync
The significance of packaging scenarios as Agent Skills is that all invocations share the same input contract, evidence thresholds, and delivery methods. They can be repeatedly reused by different Agents and combined with different collection tasks. Changing regulatory sources, regions, or industry standards doesn't require starting from scratch — the enterprise profile is passed in as a parameter, and only the collection task and domain-specific applicability rules need to be swapped.
The current entry point is financial regulatory announcements, administrative penalties, and local regulations, with future expansion possible into data and cybersecurity, consumer protection, industry-specific policies, and more. Regardless of where it expands, the core logic remains unchanged: deeper dynamic data acquisition, clearer applicability judgment, more precise risk extraction, and more actionable measures.
Conclusion
The value of this case lies not in technical showmanship but in demonstrating a pragmatic AI implementation paradigm — not letting the LLM "freestyle," but constraining it within trustworthy boundaries through rule frameworks and separation of responsibilities. MCP handles continuous dynamic data acquisition, the Skill turns data into enterprise decisions, and ultimately professional judgment, evidence trails, and delivery are placed within a single closed loop. For regulation-intensive industries, this may hold more production value than simple "AI Q&A."
Related articles

Laptops: The Last Bastion of Plaintext Secrets
Developer laptops are the last security blind spot for plaintext secrets. This article analyzes risks in .env files, shell history, and tool configs, offering practical solutions like OS keystores, dynamic injection, and short-lived credentials.

Go Microservices in Practice: Detailed Architecture for E-Commerce, AI Agent, and IM System Integration
Deep dive into integrating e-commerce, AI Agent, and IM systems under Go microservices architecture, covering unified auth, gRPC, componentized Agent engines, and group chat bots.

X Platform's Recommendation Algorithm Caught Filtering Brazilian Election Content, Reigniting Algorithm Transparency Debate
X (formerly Twitter) was found filtering Brazilian election content in its For You feed, sparking debate over algorithm transparency and free speech.