Vindle optimizes massive CSV imports by replacing PHP with Rust
Vindle reduced processing time for 27 GB of product data by moving CSV filtering from PHP to a custom Rust tool, eliminating redundant parsing passes.
Vindle, a price comparison platform, recently overhauled its data ingestion pipeline to handle a massive new integration with Bol.com. Facing a three-order-of-magnitude increase in product volume, the engineering team replaced their initial PHP-based parser with a custom Rust application. This shift allowed them to process 27 GB of gzipped CSV feeds efficiently while maintaining strict hourly update schedules.
What happened
The challenge began when Vindle integrated Bol.com, a general retailer offering over 73 million products. The data arrived in 26 gzipped CSV files totaling roughly 27 GB, with individual file sizes ranging from 1.5 MB to 4.8 GB. Since Vindle updates prices hourly, the system needed to parse, filter, and store these entries rapidly. The initial implementation used PHP’s fgetcsv() function within a Symfony application. While this approach was quick to build and sufficient for smaller feeds, it became a bottleneck when processing millions of rows where over 99% were filtered out due to irrelevant categories or missing data.
To address the performance gap, the team first introduced Xan, a command-line CSV tool written in Rust. This allowed them to offload row filtering and column selection from PHP, streaming only relevant data back into the main application. However, Vindle also required detailed statistics on why rows were rejected, such as invalid brands or unsupported shipping regions. Xan could not provide these metrics in a single pass, forcing the team to run the tool twice: once for statistics and once for filtered data. This dual-pass approach added 20–40 seconds per feed and duplicated the expensive CSV parsing work.
An attempt to parallelize these two Xan processes using Unix pipes and tee failed due to backpressure issues. The slower branch blocked the entire pipeline, and the complexity of error handling increased without delivering speed gains. Ultimately, the team wrote a small, dedicated Rust program that performed filtering and statistics collection in a single pass. This tool reads from standard input, applies exclusion rules, writes valid rows to standard output, and sends statistics to standard error. This solution restored single-pass efficiency while providing the necessary insights, all without requiring a full rewrite of the Symfony core.
Key details
- Bol.com provides over 73 million products across 26 gzipped CSV files totaling approximately 27 GB.
- The initial PHP implementation used
fgetcsv()but struggled with the volume of data requiring heavy filtering. - Filtering removes between 90% and 99.9977% of rows based on criteria like category, price validity, and shipping location.
- Using Xan for pre-processing improved speed but required two separate passes to gather both filtered data and rejection statistics.
- Parallelizing Xan processes via Unix pipes introduced backpressure bottlenecks and did not improve overall throughput.
- A custom Rust CLI tool using the
simd-csvlibrary now handles filtering and statistics in a single pass, streaming results directly to PHP.
Background
CSV parsing is often assumed to be trivial, but at scale, it becomes CPU and I/O intensive. In high-volume scenarios, the cost of reading, decompressing, and tokenizing every character in a file adds up quickly. When most rows are discarded, the overhead of parsing them fully before rejection is wasted computation. Tools like Xan leverage low-level optimizations, such as SIMD (Single Instruction, Multiple Data) instructions, to accelerate this parsing. However, generic tools may lack the specific business logic needed for complex filtering or statistical tracking. Writing a small, focused binary in a systems language like Rust allows developers to combine high-performance parsing with custom logic, avoiding the overhead of general-purpose interpreters like PHP for tight loops.
Why it matters
For teams running self-hosted applications, this case highlights the importance of identifying the right boundary between high-level application code and low-level data processing. Moving the entire ingestion workflow to Rust would have introduced significant maintenance debt and complexity. By keeping Symfony in charge of orchestration, credentials, and database persistence, the team retained the productivity benefits of PHP while solving the specific performance bottleneck with a specialized tool. This hybrid approach allows engineers to optimize critical paths without sacrificing the ease of development provided by their primary framework.
Additionally, the shift toward streaming data reduces disk dependency. The original process required writing large temporary files to disk, consuming significant storage space and I/O bandwidth. The new Rust tool processes data from standard input to standard output, enabling a diskless workflow. This is crucial for environments with limited storage or where I/O contention can impact other services. It demonstrates that performance improvements often come from architectural changes, such as eliminating redundant passes and reducing intermediate storage, rather than just optimizing code syntax.
What you can do
- Profile your data ingestion pipelines to identify if parsing or filtering is the primary bottleneck before rewriting code.
- Consider offloading heavy row-by-row processing to external binaries written in systems languages like Rust or Go.
- Avoid parallelizing identical parsing tasks via Unix pipes if the branches have unequal workloads, as backpressure will limit throughput.
- Design CLI tools to accept input from standard input and write to standard output to enable streaming and reduce disk usage.
- Keep business logic and state management in your primary application framework, using external tools only for deterministic, high-volume transformations.
- Measure the impact of multiple passes over large files; combining filtering and statistics into a single pass often yields better results than parallelization.



