Scorchsoft
From prototype to business process

AI agent harness development

Make the AI workflow that works for you work for your whole team. We build the process, tools, data access and controls so people can review actions and take over when needed.

An AI agent surrounded by relevant evidence, permitted tools, workflow rules and human decisions.

What is an AI harness?

An AI harness is the software and configuration around an AI model that lets it do a defined job. It supplies instructions, relevant information and tools, then manages execution, validation, approvals and the record of what happened.

You might already have an agent that performs useful work when you supervise it. For operational leaders, product owners and technical teams moving beyond that prototype, a harness helps make the process available to others, run it automatically where appropriate and handle exceptions without depending on one person's memory.

The outcome could be a reviewed customer response, a completed research pack or a document ready for approval. We start with that outcome and design the smallest system that can deliver it with appropriate controls.

When your AI workflow needs more structure

  • One person holds the process together by choosing inputs, correcting prompts and checking every result.
  • Several systems or stages must work together, with clear rules about what happens next.
  • The AI needs selected business data, rather than unrestricted access to everything.
  • You need approvals, traceable decisions, reliable recovery or limits on cost and execution.
  • A shared portal, scheduled job or existing application needs to trigger and track the work.

A personal agent environment or workflow builder may already meet your needs. We assess that first, then use bespoke code where it provides a useful capability, integration or control.

Design the rules around the work

Evidence becomes a draft, then validation and approval precede an external action; exceptions return for resolution.

Structure the process and the tools

We map the trigger, stages, outputs and exceptions, then decide which steps need AI judgement and which should use ordinary code. Each stage receives a defined task and an appropriate set of tools.

For example, a document-review workflow might collect files, check completeness, use AI to identify issues and create a review pack. A missing required document becomes an exception, rather than a gap the model tries to fill. Narrow tools can let an agent create a draft without giving it permission to publish.

Put controls around consequential actions

We design validation, permissions and human approval around the particular actions that need protection. A reviewer can see the proposed change, its supporting evidence and unresolved questions before it reaches another system.

Controls can include schema checks, business-rule validation, scoped credentials, approval states, execution budgets and audit records. Approval can be bound to the exact proposed action and invalidated when that action changes. Model-based checks support review, while application rules enforce boundaries such as permitted records and authorised actions. We test those controls against representative cases.

Structured queries, semantic retrieval and access filters supplying relevant evidence to an AI stage.

Give each stage the knowledge it needs

We design access to databases, documents, APIs and MCP integrations, with retrieval tailored to the question. Exact facts can come from structured queries; similarity searches can use vector embeddings over selected material.

Access filters, source references, document versions and refresh rules help the agent work from relevant evidence. We also consider what enters model calls, search indexes and logs, so context design supports performance and privacy together.

Build with Python, TypeScript and the OpenAI Agents SDK

We can build coded workflows using the Python OpenAI Agents SDK or its TypeScript counterpart, selecting the approach around your existing application and team. The SDK provides agent runtime features; we build the business process and application controls around them.

Your harness can combine function tools with MCP connectors, APIs and existing services. Using code does not require replacing working integrations. It lets us specify how those integrations participate in the workflow, what authority they receive and what must happen before a write is allowed.

We choose deliberately between a fixed sequence, an agent deciding its next step, handoffs and specialist agents used as tools. More agents are useful when they improve context, separation of responsibilities or measured performance. They also introduce cost and coordination, so a simpler design may be preferable.

Balance performance, cost and oversight

Measure the result that matters

We agree representative test cases and success criteria before optimising. These may include factual accuracy, correct tool use, appropriate escalation, completion time and review effort. We also compare process volume, current time and failure costs with the cost of building and maintaining the workflow, so the business case guides the scope.

We compare model choices and reasoning settings by stage, control context size and set limits for calls, retries and execution time. The useful measure is the full cost per successfully completed task, including external services, infrastructure and human review.

Give shared work an operational home

A portal can provide work queues, role-based access, evidence views, approvals and reporting. Behind it, hosted workers can process jobs, persist progress and connect to business systems. We plan recovery and duplicate prevention for important external actions.

The infrastructure may use AWS or Azure services such as object storage, queues, container workers and document processing. A governed data lake can help where the workflow spans many sources or historical data, but a smaller database and retrieval service may be sufficient. We identify who handles exceptions and who can pause new work if an integration or model behaves unexpectedly.

What we deliver

The scope depends on your workflow. A project can include:

  • A process map, success criteria and a specification of authority, approvals and exceptions.
  • Agent instructions, tool definitions, structured outputs and integration code.
  • Retrieval pipelines, vector-search design and access controls where needed.
  • Approval interfaces, job state, recovery behaviour and operational records.
  • A representative evaluation set and results used to compare configurations.
  • Deployment configuration, monitoring, documentation and an agreed support approach.

We agree the deliverables during planning, including code ownership, provider dependencies, data preparation, acceptance decisions, credentials, handover and ongoing operation. Where appropriate, we can extend an existing portal or application rather than create another standalone tool. Planning produces a defined first workflow, control and approval map, integration requirements, evaluation criteria and an implementation estimate against the agreed scope.

How an AI harness project works

  1. Understand the current process

    Review the prototype, human interventions, data sources and business outcome.

  2. Define the first workflow

    Agree boundaries, approval points, representative cases and measures of value.

  3. Build the harness and integrations

    Implement agent stages, tools, retrieval and the surrounding application.

  4. Evaluate with people

    Test normal and difficult cases, inspect failure behaviour and refine cost and quality.

  5. Deploy and improve

    Roll out with appropriate oversight, monitoring, documentation and support responsibilities.

Foundations you may already be able to use

OpsUPLOOP shows research, drafting, editing and human review working around shared operational records. Its outreach workflow keeps source references and approval before letters are posted; its controlled MCP tools let compatible assistants work within business permissions. We can assess its modules and integrations as a foundation for an additional workflow.

Lake On Rails records dataset ownership, classification and retention, with signed-off procedures and an audit trail. It provides the data operating model where ownership and governance need attention, alongside the underlying data platform and retrieval services your harness uses.

We can also help with AI-ready data ecosystems, business automation and application hosting and support. These components matter when they support your particular workflow; you do not need every one to begin.

For the reasoning behind this approach, read our guide to why we keep reaching for code when building AI workflows.

What Our Clients Say

  • Scorchsoft helped us take our idea for an app and make it a reality. Everything from the planning meeting to decide what we really needed to the project management and execution was great. It was delivered on time - early in fact - and on budget. Highly recommend.

    Dragonfly Intelligence logoRebecca PalserDragonfly Intelligence
  • Scorchsoft is a brilliant company with fantastic knowledge of the mobile app industry. From the project management to the development team, they have been the perfect candidate for our project, and we can't thank them enough!

    Gapped Online logoLance ChorltonGapped Online
  • We're really pleased with the work Scorchsoft has done in developing our web portal! They have been accurate with timelines and budget, delivering a solid product that allows us to monitor and manage patients remotely while they use our novel medical device at-home. The "plan - design - build" approach has worked well and saved us time in the long-run by catching requirements and issues early.

    SensTrain logoDaniel GreenSensTrain
  • I'm really pleased with how my app came out, it was exactly what I was looking for. The team at Scorchsoft are great at what they do and made the whole process as simple and easy as possible. Being someone who is not very tech savvy the set up and back end operations were done in a great easy to use manner even for myself which makes using my app stress free. Thanks to all the team!

    Mosaic Masterpieces logoRuben CarrollMosaic Masterpieces
  • The new Flourish Education website has already removed a lot of manual processes, freeing up both schools, candidates and internal employees time. We are delighted with the look and feel which is clean, professional and more engaging. We are also pleased with the decision to have an HD video background on the homepage, and building immediate trust with our clients by giving them a taste of what it looks like in the Flourish Education offices.

    James HancocksMarketing Manager, Flourish Education
  • I can't believe how quickly we started to see results with this project. Scorchsoft provided us with graphic-designed mockups of how the app would look once built, and we were able to sell the product for use by our first customer before the product was finished. Since launching in March, we have secured a major television network as a client who now uses Image Approvals to manage the talent approval process for their productions.

    Aimee SpinksMD, ImageApprovals.com

Start with the workflow you already want to improve

Show us one workflow, the tools it uses and where you currently intervene. In an initial discussion, we can explore the controls and integrations it needs, whether your current platform can provide them, and a sensible next step. Detailed scoping and deliverables follow through an agreed planning engagement.

AI harness development FAQs

The model generates or interprets information. An agent uses a model with instructions and tools to perform work. The harness supplies the surrounding execution process, context and controls. Terminology varies, so we define the components and responsibilities explicitly during planning.

That can work well for supervised tasks and experimentation. A shared or unattended business process may need explicit job state, narrow permissions, approval screens, recovery and integration with your application. We assess the controls available in your current platform before recommending bespoke development.

Yes. They can provide instructions, examples and reusable tools. We identify what should remain model guidance and what needs a schema, permission, application rule or mandatory approval. The prototype gives us a useful starting point, while testing establishes what is suitable for wider rollout.

Yes. We can use the OpenAI Agents SDK in Python or TypeScript and connect it to your application. We select the language around the workload, existing code and maintenance needs. The SDK does not replace the surrounding deployment, data engineering or business controls.

AI judgements remain probabilistic. A harness can make important surrounding rules explicit and enforce permissions, approval and validation. We test the overall workflow, include escalation for insufficient evidence and keep appropriate human oversight for consequential decisions.

We measure representative jobs, compare models by stage, limit calls and retries, and design focused context and tool responses. We include external-service fees, hosting and review effort when assessing the cost per completed outcome. Savings depend on the task and measured results.

We can design retrieval with access filtering and narrow tool scopes. We review source systems, indexes, model calls, logs and retention together. Hosting the harness yourself does not automatically keep inference inside your infrastructure if it calls a hosted model provider.

A portal helps when a team needs queues, permissions, shared records and approvals. A data lake may help with many sources, historical data and governance. Some workflows need neither. We recommend the foundations that fit the process rather than require a large platform from the start.

It runs in an agreed application or cloud environment, often as an API and background workers. We can plan AWS or Azure infrastructure with queues, object storage, secrets and monitoring. Execution duration, volume and data requirements determine the appropriate arrangement.

Scope depends on integrations, data readiness, approval interfaces, evaluation requirements and the surrounding application. We define a bounded first release during planning and provide an estimate against that scope. A prototype helps establish feasibility, but does not by itself establish the effort required for deployment.

We agree monitoring, support, ownership and a process for reviewing changes to models, instructions and integrations. Evaluation cases help check that an improvement in one area has not damaged another. Operational feedback can guide further development and changes to oversight.

Need help building your ideas?

Tell us where you're headed and we'll come back with our thoughts, a realistic plan and a rough cost estimate. Scorchsoft is a UK-based team of app, portal and AI developers, working in-house from Birmingham's Jewellery Quarter.