Portus gateway swaps Pingora for Rama engine to cut memory use
Portus version 0.2.4 adopts the Rama network stack, delivering up to 31 percent higher throughput and significantly lower memory usage compared to its previous Pingora base.
Bài này hiện chỉ có bản tiếng Anh.
The Portus API and AI gateway project has released version 0.2.4, marking a significant shift in its underlying network infrastructure. By replacing the Cloudflare-built Pingora framework with the Rama engine, the developers report substantial improvements in resource efficiency and request handling speed. This change affects how the gateway processes traffic, manages memory, and scales within Kubernetes or standalone environments.
What happened
Portus is an open-source gateway designed to manage API traffic and artificial intelligence model interactions. In its earlier releases, the project relied on Pingora 0.9, a high-performance proxy framework developed by Cloudflare. However, after conducting side-by-side benchmarks on identical hardware, the Portus team determined that the Rama 0.4 engine offered superior performance characteristics for their specific architecture. Consequently, version 0.2.4 makes Rama the default network stack.
The decision was driven by empirical data collected on September 15, 2026. The tests showed that Rama delivered between three and thirty-one percent higher throughput across various payload sizes when compared to Pingora. More critically for operators managing large-scale deployments, Rama consumed two to two-point-six times less memory. These metrics convinced the maintainers to switch the default engine, although Pingora remains supported for users who may have specific legacy requirements.
Beyond the network stack swap, the release reinforces the gateway’s architectural separation between control logic and data processing. The system uses a shared store to compile configuration changes into a protobuf format, which is then distributed to data planes via mTLS gRPC. This ensures that configuration updates are applied atomically without dropping active connections, a critical feature for maintaining uptime during rolling updates or policy changes.
Key details
- Default engine change: Version 0.2.4 switches the default network stack from Pingora 0.9 to Rama 0.4.
- Performance gains: Rama demonstrates a three to thirty-one percent increase in throughput across all tested payload sizes.
- Memory efficiency: The new stack uses two to two-point-six times less memory than the previous implementation.
- Atomic reloads: Configuration updates arrive over secure gRPC channels and swap without interrupting existing connections.
- Standalone mode: The gateway can operate outside Kubernetes using a single YAML file, with built-in ACME certificate management.
- AI ledger isolation: Usage tracking and key management run on a separate ledger service that does not sit on the critical request path.
Background
To understand the significance of this change, it helps to distinguish between the control plane and the data plane in modern gateways. The control plane manages configuration, such as routing rules and security policies, while the data plane handles the actual network traffic. Portus decouples these layers, allowing the core logic to remain independent of the underlying network framework. This modularity enabled the team to test different engines like Pingora and Rama without rewriting the entire application.
Rama and Pingora are both Rust-based frameworks known for high performance, but they optimize for different trade-offs. Pingora is battle-tested in massive global networks, while Rama is often praised for its flexibility and low-level control. By abstracting the network stack, Portus can leverage the specific strengths of each engine. In this case, the lightweight nature of Rama aligned better with the gateway’s need for efficient memory usage and fast local lookups for authentication and rate limiting.
Why it matters
For teams running self-hosted infrastructure, memory efficiency directly translates to cost savings and stability. A reduction in memory usage by a factor of two or more means that organizations can run more gateway instances on the same hardware or downsize their virtual machines without sacrificing performance. This is particularly relevant for AI gateways, which often handle large payloads and require rapid token budget checks. Lower memory overhead reduces the risk of out-of-memory crashes during traffic spikes.
The atomic configuration reload mechanism also addresses a common pain point in DevOps: downtime during updates. Traditional proxies often require restarts or brief interruptions when applying new TLS certificates or routing rules. Portus’s approach ensures that connections are never dropped during these changes. For IT leads managing service level agreements, this reliability minimizes the operational burden of maintenance windows and reduces the risk of user-facing errors during routine configuration adjustments.
Furthermore, the ability to run the gateway in standalone mode using a simple YAML file lowers the barrier to entry for smaller teams. Not every organization needs a full Kubernetes cluster to benefit from advanced API management features like automatic Let's Encrypt certificate renewal. This flexibility allows developers to deploy consistent routing and security policies across diverse environments, from local Docker Compose setups to production virtual machines.
What you can do
- Benchmark your current setup: If you are using Portus 0.2.3 or earlier, test the upgrade to 0.2.4 in a staging environment to verify the memory savings in your specific workload.
- Review standalone configurations: Evaluate if any non-Kubernetes services could benefit from the standalone YAML mode, especially for edge cases where a full cluster is overkill.
- Monitor connection drains: Observe how the SIGTERM drain behavior affects your load balancers during deployments to ensure smooth traffic shifting.
- Check AI ledger settings: If you use the AI gateway features, review the opt-in ledger configuration to ensure token budgets and key fetching align with your security policies.
- Validate TLS automation: Test the built-in ACME client to confirm that certificate renewals at two-thirds of their lifetime proceed without manual intervention.
- Plan for multi-round testing: Note that the current benchmarks are from a single round; prepare to review the upcoming three-round comparison data for more robust performance insights.



