Start with a workflow

A useful starting point is a recurring task with a clear result: finding information in approved documents, preparing a reply draft or sorting incoming requests. Describe the current process, the people involved and the effort it takes. This reveals where AI could help and where straightforward automation would be sufficient.

Prioritise expected benefit, technical feasibility and the consequences of errors. Frequency alone is not enough: a rare but consequential mistake can outweigh an apparent time saving. A bounded use case with available subject expertise is easier to evaluate than an assistant expected to answer every question about the organisation immediately [1].

Clarify data and access permissions

Decide which data may be processed, for what purpose and who may see the result. Include personal data, business secrets and retention periods. Involve those responsible for data protection and, where applicable, employee representatives early [2].

Document retrieval must respect the current user's permissions when supplying context to AI. A shared search index must not bypass existing access controls. Review the entire processing chain, including document preparation, model calls, logging and external services.

Select the tool, model and operating approach

Compare existing enterprise software, model APIs and a custom integration using the same tasks. Evaluate output quality, connections to existing systems, access management, costs and the ability to change providers. A single benchmark does not settle these questions.

Hosted services, European providers and local operation offer different options. A company's headquarters do not guarantee a processing location; regional cloud offerings also depend on specific settings [3]. Our comparison of internal AI assistants examines this decision. Agents with tool access become relevant when the workflow requires actions to be carried out autonomously.

Evaluate the pilot against agreed criteria

Establish a baseline before starting: processing time, error rates and rework in the current process. Then test with approved, representative tasks, including difficult cases and missing information. Measure total handling time, including checking and correction, rather than only the speed of the AI response.

Define minimum quality, a cost budget and stopping criteria. Retain human approval for consequential outputs. Tool access also needs narrowly scoped permissions and technical checks; a good system prompt does not replace these controls [4]. Securing MCP covers the interface side.

Involve the people who will use it

Workshops should centre on real tasks: writing appropriate inputs, checking outputs, recognising limitations and reporting errors. Record which tools are approved for which data and where people can get support. Pilot feedback identifies missing guidance, process changes and additional training needed before wider adoption.

Assign responsibility for ongoing operation

Name owners for business value, data and technical operation. Document approved use cases, permissions and model versions. Re-evaluate changes to models, data sources and integrations against the pilot's test set [1].

Ongoing operation includes quality checks, cost monitoring, an incident reporting route and a workable fallback process. Logs should help explain failures without unnecessarily storing sensitive inputs. Also decide when a use case should be adjusted, paused or retired.

Decision point

Expand adoption when benefit and quality are demonstrated, access controls and approvals work, and operation and support can be sustained. Otherwise, improve the pilot in specific areas or stop it.

Sources

  1. AI Risk Management Framework 1.0 NIST, 2023
  2. Orientierungshilfe Künstliche Intelligenz und Datenschutz Datenschutzkonferenz, 2024
  3. Data, privacy, and security for Foundry Models sold by Azure Microsoft Learn
  4. LLM06:2025 Excessive Agency OWASP GenAI Security Project

Read on