Skip to content
6 min read

Rails for AI agents: why we build on it and what to demand of any stack

Why we build agentic software on Rails, and a portable checklist for any stack before you let an agent near production data.

Mauricio Zaffari

Rails for AI agents: why we build on it and what to demand of any stack

In a hypothetical example, a company hands a ticket to a coding agent. The agent has to work out where the code lives, which rules the business follows, how to run the tests, and where the credentials sit. In a stack with no conventions, every one of those questions has a different answer per project, and the agent burns much of its effort guessing what a person would know from experience.

At Develoz we build our solutions on Rails, and this post explains why. You get two things: the criteria we use to judge a stack for agentic work, and a portable checklist to demand of any stack before an agent touches production data.

Why do conventions matter more now?

An agent learns from the code that exists. When a framework decides where things go, the agent stops searching and starts executing. The official Rails AI page puts it in one line: "Less code means more context." That is literal: every line the framework removes is one fewer line the agent has to read, interpret, and get wrong.

The same principle holds at the scale of training data. Rails creator David Heinemeier Hansson argues on that page that convention over configuration "set the path for 20+ years of great training data for AI to use today." The framework's patterns are in the data the models learned from, which cuts down guesswork at run time.

What do the public benchmarks show?

That reading is not ours alone. The Rails Foundation commissioned Evil Martians to benchmark AI models working on Rails codebases: the "Agents on Rails" benchmark. The tasks are open at github.com/rails/ai-evals and graded on behavior: a hand-rolled fix passes just like an idiomatic one, and what counts is the observed outcome.

The benchmark's numbers are honest about limits. On small, well-scoped tasks, the best model solved 92% of cases. On complete feature tickets, the best result dropped to 35%, rising to 53% at maximum effort. For a buyer: scope weighs more than model choice.

ScenarioCited result
Small, well-scoped tasks (stage 1)92% solved by the best model
Complete feature tickets (stage 2)35% for the best result
Complete tickets at maximum effort53%
Complete tickets after the credential lockdown12% and 17%

A later report matters more than the wins. One of the models found the test harness's own API key and started searching the web for answers. After lockdown the results fell to 12% and 17%, and the foundation rechecked all 2,300 runs. The takeaway is direct: an agent tends to explore whatever it can reach, and the credential boundary is our responsibility, not the model's. This is the same argument we make in human review and automation limits.

flowchart TD
  A[Task handed to the agent] --> B[Read project conventions and rules]
  B --> C[Generate code]
  C --> D{Framework tests}
  D -- fail --> B
  D -- pass --> E[Credential and permission boundary]
  E -- blocked --> F[Human review]
  E -- allowed --> G[Merge]

What does the stack give you out of the box?

Rails 8, announced November 7, 2024, shipped Solid Queue, Solid Cache, and Solid Cable inside the framework, making Redis optional, and made SQLite a viable production database. Fewer moving parts means less context for the agent to assemble and less infrastructure for it to break.

The strongest point is the built-in verification loop. Testing is part of the framework, so any Rails environment already has, with no configuration, the mechanism an agent needs to check its own work. The Pragmatic Engineer, on April 8, 2026, called Rails "one of the most token-efficient ways of building web apps and well-suited for agent workflows", citing the fact that testing is part of the framework. It is the same criterion we apply in our gates for qualifying a process for AI.

Did the ecosystem keep up?

The RubyGems guides are now also published in formats models read directly, and tools exist that let an agent inspect a project's routes, models, and schema instead of inferring them. When the stack describes itself, the agent reads instead of guessing.

The same value shows up in integration with the rest of the operation. An agent that touches the business has to talk to the systems already running, which is where our post on integrating AI without replacing your ERP picks up: the stack is one part; the connection to what already runs is the other criterion.

Where does Rails still fall short?

Dynamic typing gives the agent fewer automatic checks, so wrong method names can slip through, and part of the industry has moved the other way, toward typed languages. GitHub's Octoverse 2025 reported TypeScript passing both Python and JavaScript in contributor count. Ruby's AI library ecosystem is also smaller and less mature than Python's. And 37signals announced that HEY's backend is being rewritten toward Rust; that work is underway, not finished.

We weighed this and stayed with Rails, partly with tooling of our own. The rails-quality-assurance gem, which we develop and maintain, moves outside the runtime the checking that dynamic typing does not provide: RuboCop, Reek, and Flay for style, code smells, and code duplication, Brakeman and dependency auditing, RSpec with a 100% minimum coverage default, and a CI pipeline with a pre-commit hook. The agent runs that battery as part of its own verification loop. A dedicated post on the gem is coming.

In summary

The argument is not that Rails is the future. It is that conventions, lean code, tests in the framework, and machine-readable documentation solve precisely the problems an agent has. From that comes the checklist we demand of any stack before an agent touches production data:

  • Conventions the models already know: project patterns match what the models learned, and documentation exists in a machine-readable format.
  • A built-in verification loop: the agent can run tests and know on its own whether the work passes.
  • Few moving parts: the less infrastructure outside the framework, the less context the agent has to assemble.
  • Introspection tooling: the agent reads routes, schema, and rules from the stack itself, instead of inferring them.
  • Credential boundaries set before the agent ships: the benchmark showed what happens when something sensitive stays reachable.
  • Human review where the obligation exists: the approval point depends on where the process creates a financial or business risk.

This checklist applies to Rails, to any language, and to any framework. If your stack passes it, you have the basic conditions in place. If it does not, fix the boundaries before you scale the agent's access.

Want to assess where agents and AI could fit into your operation, connected to the systems you already run? Request a diagnostic.