Fireworks AI recently launched FireRouter with Opus, the first cache-aware router in the market optimized for the Claude Opus series and now available as an independent routing model through a serverless endpoint. After more than a month of internal A/B testing, this solution achieved an accuracy rate of 98.1% for encoding tasks compared to using Opus alone, while reducing costs by 57%.

FireRouter's core logic is to evaluate the suitability of each model in the model set for the current task during each user turn, estimate the processing cost of each model (including prompt caching cost), weigh whether the savings from switching models are worth the loss of cached content, and finally route to the model that achieves the best balance between quality and cost. The current routing pool includes Claude Opus5.5, GLM5.3, and GLM5.3Flash, and will continue to be adjusted as new models are released.

image.png

The test data comes from internal encoding traffic, with sessions randomly assigned to either the FireRouter with Opus group or the control group that only uses Opus, with the same workload and users. The results show that the cost per session dropped from $15.36 to $6.63, a reduction of 57% (±19 percentage points, 95% confidence interval). In terms of accuracy, FireRouter with Opus scored 78.7% in scoring rounds, while using Opus alone was 80.2%, a difference of just 1.5 percentage points (±1.3), reaching 98.1% of Opus's overall accuracy.

In terms of cache hit rate, FireRouter with Opus reached 94.2%, while using Opus alone reached 97.8%, a difference of 3.6 percentage points. Fireworks AI explained that this is a small but intentional trade-off — since open-source models can handle many routine turns in internal encoding work, cache-aware routing significantly reduces overall costs by directing simple tasks to cheaper models.

image.png