Responsibilities
- Manage and scale PostgreSQL and Supabase databases with hybrid search capabilities, focusing on schema design, data modeling, and performance optimization as data volume grows indefinitely.
- Implement end-to-end data quality processes including validation checks for third-party and vendor data, deduplication, entity resolution, data provenance tracking, and continuous monitoring.
- Build and maintain a communications data layer that securely stores raw emails and call transcripts, links them to relevant people and organizations, and enables efficient search.
- Develop scalable ingestion and enrichment pipelines for funding data, market news, and contact/company intelligence, optimized for cost-efficiency and data freshness.
- Design and evolve a knowledge graph within Postgres using node and edge tables to represent companies, investors, funding rounds, and news, with provenance tracked for each fact.
- Construct a unified, reliable data backbone that serves as the single source of truth for all campaigns, AI agents, and product features.
Benefits
- Fully remote and asynchronous work environment.
- Partial overlap with US Eastern time zone required, but not full-time US hours.
- Meetings are grouped on Mondays and Thursdays to preserve uninterrupted focus time the rest of the week.
- Access to premium AI development tools including paid subscriptions to Claude Code, Cursor, and leading AI models.
- Collaborate directly with the GTM lead and founding engineering team, where your data systems directly enable all product and campaign development.
Work Arrangement
Remote (Worldwide)
Responsibilities
- The database itself: Postgres and Supabase with hybrid search, schema design, modeling, scaling, and performance as it grows without a ceiling. The agents that read it run on Pydantic AI and the Claude Agent SDK, on AWS. We are consolidating into pgvector, not buying a vector DB.
- Data quality end to end: validation gates for vendor and third-party data, dedup, entity resolution, provenance, monitoring.
- The communications layer: raw emails and call transcripts stored, linked to the right people and companies, and searchable.
- Ingestion and enrichment pipelines: funding rounds, market news, and contact and company research at scale, engineered for cost and freshness.
- The knowledge graph: companies, investors, funding rounds, and news as entities and relationships - node and edge tables in Postgres, provenance on every fact.
- The unified data layer: one clean spine that every campaign, agent, and product feature reads from.
Benefits
- Fully remote and async.
- Your day overlaps with US Eastern time for a few hours - not full US hours.
- Meetings batch on Mondays and Thursdays, the rest is deep work.
- The best AI tooling, paid (Claude Code, Cursor, top models).
- You work alongside our GTM lead and our founding engineers, and your layer feeds everything they build.