Machina Arena · Evaluate
AvailableCompare agents on real tasks.
Run the same task across agents, workflows, or model configurations. Review inputs, outputs, and criteria before deciding what changes.
- 01Define. Choose the task set, expected behavior, and evaluation criteria.
- 02Run. Evaluate comparable configurations against the same material.
- 03Review. Inspect strengths, failures, and tradeoffs rather than reducing quality to one opaque score.
- 04Decide. A responsible owner selects the configuration or the next change to test.
The job
What Machina Arena does
Arena gives teams a structured way to define relevant tasks and criteria, compare configurations, and use the resulting evidence to inform a human decision.
- 01
Task-based evaluations
Test against examples and conditions that represent the real workflow.
- 02
Configuration comparison
Compare agents, workflows, or model configurations on the same criteria.
- 03
Inspectable evidence
Keep evaluation inputs, outputs, and criteria visible to the people making the decision.
- 04
Human decisions
Use results to guide the responsible owner's next configuration change.
What can Arena compare?
Arena can compare agents, workflows, or model configurations against tasks and criteria chosen for the use case.
Who chooses the configuration?
A responsible owner reviews Arena's evidence and decides which configuration to use or test next.
Machina Arena in the stack
How Arena connects to the wider platform.
Choose a flow, then select a component to see how the work moves.
Use the component buttons to explore the architecture.
Select a component
Evaluate flow
Compare. Learn. Decide.
Arena compares candidates against useful test cases and criteria. A responsible owner reviews the evidence and approves any configuration change.
Selected component
Machina Arena
Arena compares candidates against relevant test cases and criteria, keeping execution evidence available for review.
Explore Machina Arena →View connections
- Context → Arena: Test cases and ground truth
- Studio → Arena: Candidates and execution evidence
- Arena → Review: Comparison evidence
Evaluation example
Comparing recap configurations before a new competition
An editorial team evaluates factual grounding, terminology, completeness, and review effort across candidate configurations.
Inputs
- Representative fixtures
- Expected facts and terminology
- Editorial evaluation criteria
Output
A comparison record that supports a human configuration decision.
One platform
