Data & spreadsheets

Reducing LLM costs by moving spreadsheet aggregation to the server side

A data engineer reduced token usage by 2,300 times for large spreadsheets by performing calculations on a local server instead of sending raw rows to an AI model.

Excel Import Mapper preview

Data engineer Andrei Kushniarou developed a local Model Context Protocol (MCP) server called gsheets-mcp to handle large Google Sheets datasets without compromising privacy or incurring excessive AI costs. Published in late September 2026, his technical deep-dive demonstrates how shifting data processing from the language model to the server can reduce token consumption from millions to under a thousand for complex financial queries.

What happened

Kushniarou needed to analyze e-commerce settlement reports containing tens of thousands of rows using Claude Opus 5.5. He rejected third-party MCP servers due to data privacy concerns and built a local solution that interacts directly with the Google Sheets API. To test the system, he used a synthetic Amazon Settlement report with 50,009 rows and 24 columns, asking the model to break down revenue by amount type. The initial approach of sending raw data to the model proved inefficient and expensive, prompting a series of optimizations.

The first version of the server returned data as indented JSON, resulting in 1.83 million tokens for a single query. This exceeded the context window limits of most models and triggered client-side errors in Claude Code, which caps tool call outputs at 25,000 tokens. Additionally, the cost of input tokens for Opus 5.5 is $4 per million, making repeated reads of large datasets financially unsustainable. Kushniarou noted that even if the data fit, relying on an LLM for arithmetic across 50,000 rows is unreliable for accounting purposes.

Through iterative improvements, Kushniarou reduced the token count to just 787 for the same query. The final architecture uses server-side aggregation, where the local server performs grouping and summation before sending only the results to the model. This approach ensures that sensitive financial data remains within the user's infrastructure while providing the AI with concise, accurate summaries rather than raw, noisy datasets.

Key details

  • Initial inefficiency: The v0.1 server sent indented JSON records, costing 1.83 million tokens per query and exceeding client output limits.
  • Format optimization: Switching from indented JSON to Tab-Separated Values (TSV) reduced token usage per row from 199 to 116, a 42% decrease.
  • Column selection: Allowing the model to request specific columns rather than entire rows reduced total tokens from 1.03 million to 580,000.
  • Server-side aggregation: The gsheets_aggregate function performs sum, count, and group operations locally, reducing the final token count to 787.
  • Cost reduction: The optimization lowered the cost per query from approximately $23 to $0.39, a reduction of over 98%.
  • Privacy preservation: All data processing occurs on a local server, ensuring that raw financial records are never transmitted to external AI providers.

Background

Model Context Protocol (MCP) is an open standard that allows AI assistants to connect securely to local data sources and tools. In this context, an MCP server acts as a bridge between the AI model and Google Sheets. Instead of the model trying to interpret a massive spreadsheet directly, it sends instructions to the MCP server, which retrieves and processes the data. Tokens are the basic units of text that language models process; both input and output text are broken into tokens. Complex data formats like JSON with indentation consume significantly more tokens than compact formats like CSV because they repeat keys and include whitespace characters.

Why it matters

For teams running their own software or managing internal data pipelines, this case study highlights the hidden costs of naive AI integration. Sending raw database dumps or large spreadsheets to an LLM is not only expensive but often technically impossible due to context window limits. By moving computation closer to the data, organizations can leverage AI for high-level reasoning while relying on traditional, reliable code for data manipulation and arithmetic. This hybrid approach prevents the "hallucination" of mathematical results and ensures data integrity.

Furthermore, the privacy implications are significant for small and mid-sized companies. Many businesses hesitate to adopt AI tools because they fear exposing customer data or financial records to third-party APIs. A local MCP server architecture allows these companies to use powerful AI models for analysis without ever letting sensitive raw data leave their controlled environment. This setup complies with strict data governance policies while still enabling advanced automation and insight generation.

What you can do

  • Audit your data formats: Check if your AI tools are receiving indented JSON or verbose XML. Switch to compact formats like CSV or TSV to reduce token usage immediately.
  • Implement column filtering: Configure your data connectors to allow the AI to request only the specific columns needed for a task, rather than downloading entire tables.
  • Move aggregation to the server: Use SQL or server-side scripts to perform sums, counts, and averages before passing data to the LLM. Never ask an LLM to sum thousands of rows.
  • Use local MCP servers: For sensitive data, deploy local MCP servers that interact with your internal databases or spreadsheets, keeping raw data within your network.
  • Validate data types: Ensure your data pipeline handles mixed formats, such as currency symbols or percentages, correctly before aggregation to avoid silent errors.
  • Monitor token usage: Track the token count for each tool call to identify inefficiencies. Aim to keep tool results under the client’s output limit to avoid truncation errors.

More news

All news