Self-hosting

Sifthound offers a self-hosted, drop-in replacement for Tavily search API

GitHub user khsarvar released Sifthound, an open-source alternative to Tavily that runs on your own infrastructure using SearXNG and supports MCP.

Illustration of a self-hosted search server connecting to a developer workstation
Illustration created for this article

On September 25, 2026, GitHub user khsarvar published Sifthound, an open-source project designed to serve as a self-hosted alternative to the Tavily web search API. This release targets development teams building AI agents and large language model applications who prefer to manage their own infrastructure rather than rely on external software-as-a-service providers. The tool replicates the core functionality of Tavily while running entirely on local or private servers.

What happened

Sifthound is built to mimic the Tavily API structure, offering endpoints for search, content extraction, crawling, and site mapping. By matching the request and response formats of the original service, it allows developers to switch from the hosted Tavily platform to their own instance with minimal code changes. The project relies on SearXNG, a privacy-respecting metasearch engine, to gather results from public search providers like DuckDuckGo and Brave. For content processing, it uses trafilatura to extract clean text and markdown from web pages, and BM25 algorithms for relevance ranking instead of neural reranking.

The release includes support for the Model Context Protocol (MCP), enabling direct integration with AI coding assistants such as Claude Desktop and Cursor. Developers can run Sifthound via Docker Compose, which bundles the necessary SearXNG instance, or install it directly using pip. The project is licensed under the MIT License, permitting free use and modification, though the name itself is protected by a separate trademark policy that requires forks to use different branding.

Key details

  • Drop-in compatibility: The API accepts the same parameters as Tavily for /search, /extract, /crawl, and /map endpoints, working with official Python SDKs and LangChain integrations by simply changing the base URL.
  • No search API keys required: Search queries are routed through SearXNG, eliminating the need for individual API keys from search engines, though an Anthropic key is optional for generating direct answers.
  • Security measures: The system blocks requests to private, loopback, and link-local addresses to prevent server-side request forgery, including checks against DNS rebinding attacks during redirects.
  • MCP support: It exposes an MCP server over HTTP at /mcp and supports stdio execution for local clients, allowing AI agents to use search tools directly within their workflow.
  • Ranking differences: Relevance scores are calculated using BM25 blending rather than Tavily’s neural reranker, and features like include_image_descriptions are accepted but currently ignored.
  • Error handling: If all upstream search engines rate-limit or CAPTCHA the server IP, Sifthound returns a 502 error with specific reasons rather than an empty result set, helping agents distinguish between no results and temporary failures.

Background

Tavily has become a popular choice for AI developers because it simplifies the complex task of fetching real-time web data for large language models. Normally, an agent would need to query multiple search engines, parse messy HTML, remove ads and navigation elements, and rank the results for relevance. Tavily handles this pipeline as a managed service. However, relying on a third-party API introduces dependencies on external uptime, pricing models, and data privacy policies. Self-hosting solutions aim to give teams control over this data pipeline, keeping sensitive query logs internal and avoiding per-request costs.

SearXNG acts as the backbone for Sifthound’s search capabilities. It is an open-source metasearch engine that aggregates results from various providers without tracking users. By combining SearXNG with extraction tools like trafilatura, Sifthound recreates the value proposition of Tavily: delivering clean, structured data to AI models. The inclusion of MCP support reflects a growing trend in AI development where tools are standardized to allow seamless interaction between different AI clients and local resources.

Why it matters

For teams running their own software, controlling the search layer of an AI agent reduces operational risk and cost. When you host Sifthound yourself, you are not subject to the rate limits or price hikes of a commercial API provider. This is particularly important for applications that perform high volumes of searches, where per-call fees can accumulate quickly. Additionally, keeping search queries on-premise ensures that proprietary research or internal data references do not leave your network, addressing compliance requirements for many enterprises.

The drop-in nature of the API means that engineering teams do not need to refactor their existing codebases to benefit from self-hosting. If a project already uses langchain-tavily or the official Tavily Python client, switching to Sifthound requires only a configuration change. This lowers the barrier to entry for organizations wanting to move away from SaaS dependencies without incurring significant technical debt. It also provides a fallback option if external services experience outages, ensuring business continuity for critical AI workflows.

However, self-hosting comes with trade-offs. Sifthound does not offer JavaScript rendering, meaning it cannot extract content from single-page applications that rely heavily on client-side scripting. Teams needing this feature may still require tools like Firecrawl. Furthermore, because it relies on public search engines via SearXNG, it is susceptible to the same CAPTCHA and rate-limiting issues that affect any automated scraper. Understanding these limitations helps teams decide if the trade-off between control and convenience is right for their specific use case.

What you can do

  • Test the API locally: Clone the repository and run docker compose up to start a local instance, then use curl or the Python SDK to verify search and extraction capabilities against your current workflow.
  • Configure API keys: Set the API_KEYS environment variable to restrict access to your Sifthound instance, ensuring that only authorized applications can trigger search or crawl operations.
  • Adjust SearXNG settings: Edit the docker/searxng/settings.yml file to enable more search engines, which helps mitigate rate-limiting issues when one provider temporarily blocks your IP address.
  • Integrate with MCP clients: Add Sifthound to your Claude Desktop or Cursor configuration using the provided JSON snippet to enable natural language search directly within your coding environment.
  • Monitor error logs: Watch for 502 errors indicating failing engines, and implement retry logic in your application to handle temporary blocks from upstream search providers gracefully.
  • Evaluate rendering needs: Assess whether your target websites require JavaScript rendering; if they do, plan to integrate a separate rendering service alongside Sifthound for those specific domains.

More news

All news