Artificial intelligence

How Forter Enabled Around 200 Employees to Build AI Agents in Two Weeks

Ben Maraney of Forter presents a practical experiment in scaling AI-agent development across research and development teams, using a centralized MCP server, multiple platforms for experimentation and execution, and reduced legal and security barriers. The experiment shows that tool accessibility and transparency of tool calls can sometimes matter more than adopting a complex retrieval-augmented generation architecture or advanced evaluation systems from the outset.

2026-09-28
6 min read
0 views
certi.news Editorial Team
How Forter Enabled Around 200 Employees to Build AI Agents in Two Weeks

Forter managed to involve around 200 people in building AI agents during a practical two-week sprint after designing an environment that reduced the need for programming expertise and allowed analysts and engineers to test their ideas quickly. Ben Maraney, the company’s principal engineer, presents the experiment as a lesson in removing unnecessary complexity rather than attempting to solve all the challenges of building agents at once.

The Beginning: A Practical Sprint with a Five-Week Deadline

The goal was to train research and development teams to create their own agents, including analysts from legal, psychological, and scientific backgrounds, some of whom had never written SQL. The team needed to prepare the environment within five weeks and then run a two-week building sprint. The design therefore focused on three areas: tools, platforms, and removing barriers for users.

A Centralized MCP Server Instead of Scattered Tools

Forter relied on an internal server based on the MCP protocol called Toolchain. The server provided a single interface for discovering and testing tools and creating each agent’s connections. It also allowed users to select only the tools the agent needed, rather than granting it broad access that could confuse it or lead it to use unsuitable tools.

The number of tools grew from around 20 when the team was established to nearly 60 when the sprint began, and then to almost 100 two weeks after it ended. Standardizing the repository and providing clear examples and configuration files helped add new tools quickly, while also providing governance capabilities such as monitoring token consumption and tool usage.

Instead of building a custom retrieval-augmented generation (RAG) system within a short period, the company connected Toolchain to the Glean platform used to search internal sources such as Confluence, Asana, Jira, Slack, and Salesforce. It provided three types of tools: searching for relevant documents and passages, reading an entire document, and summarizing it according to a question or topic specified by the agent.

From a No-Code Platform to Customizable Solutions

Forter used LibreChat to provide a fast, low-code conversational experience. Users could see the tools the agent called, the queries and parameters it sent, and the results it received. This helped them understand agent behavior and discover problems early, although MCP integration was sometimes unstable, with limited control over versions and customization.

For more complex cases, the company provided template repositories that gave developers full control over the code and allowed them to add tools and subagents, but they were slower to set up and deploy, particularly for analysts. Forter later created an internal interface called AI Hub that allowed users to select the model, write the system prompt, specify the tools, and share the agent in just a few steps.

For non-interactive agents, Strands was used with Argo Workflows to run them according to schedules or events such as the creation of Jira and Asana tickets. Maraney notes that having an existing scheduling and execution system may eliminate the need for new agent-specific infrastructure.

What Was Actually Built?

The use cases included agents that helped analysts formulate better hypotheses and experiments, reviewed post-incident reports, and analyzed declines in merchant performance using Snowflake and Databricks notebooks. The company also developed expert agents such as Layla and Penny to access code, configurations, and data and answer questions related to transaction and billing decisions. Layla’s use expanded to customer success and support teams after its system knowledge became broader than that of any single individual.

Examples of non-interactive agents included suggesting initial configurations for a new merchant based on similar merchants, preparing preliminary research when a support ticket arrived, and analyzing anomaly alerts and linking them to code changes. Forter also tested an incident-response agent, but discovered that its apparently excellent results depended primarily on copying root-cause analyses previously written by humans in BetterNext reports.

What Is Changing in Practice?

This experiment shows that observability is not a secondary feature. Showing the tools, parameters, and results on which the agent relied was essential to discovering that the incident-response agent was not actually inferring the root cause. Forter therefore later added Langfuse to track agent sessions, tool calls, and workflow paths.

The experiment also shows that using multiple platforms can be intentional: a no-code platform for rapid experimentation, software-based solutions for customization, and agents that operate on events or schedules. At the same time, the company encountered problems with creating a separate repository for each agent, and then returned to a smaller number of repositories containing related agents, making it easier to introduce shared capabilities such as tracing.

Maraney also cautions that general evaluation metrics such as relevance, coherence, and safety may create a false impression of quality. A useful evaluation must test the correctness of the answer, the use of appropriate tools, and the retrieval of the correct context—things that are more difficult and require precise knowledge of what a good answer means. Therefore, advanced evaluations can be postponed in internal experiments in which a human remains in the decision-making loop, but this should not be regarded as a permanent substitute for evaluation.

Scaling initiatives also need to involve legal and security teams early, particularly regarding where data is processed and retained. Maraney says that using a cloud service such as Amazon Bedrock, together with documentation of data-retention policies, helped address these concerns within the organization’s controls instead of turning every agent into a separate approval case.

News source
InfoQ - Architecture Articles
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news