Suggesting data imports

SkillDatabases & data

This skill guides an AI agent to notice when the data a person asks about lives outside PostHog and to suggest importing it through the data warehouse. It covers revenue, payments, subscriptions, billing, CRM deals, support tickets, ad spend, and production database tables, plus joining that external business data with PostHog product events.

Use Suggesting data imports in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Suggesting data imports and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Suggesting data imports skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Have an AI agent that can call PostHog tools such as posthog:external-data-sources-list and posthog:external-data-schemas-list.

Suggesting data importsStart free

What your AI can do with it

  • Spot when a question needs data PostHog does not collect natively
  • Check existing sources with posthog:external-data-sources-list
  • List available tables with posthog:external-data-schemas-list
  • Identify the right source type for SaaS tools, ad platforms, or databases
  • Guide setup for importing external data into the data warehouse
  • Recognize query failures caused by missing tables

Getting started

  1. Have an AI agent that can call PostHog tools such as posthog:external-data-sources-list and posthog:external-data-schemas-list.
  2. Add the suggesting-data-imports skill to that agent so it can recognize when a question needs external data.
  3. When a query fails or mentions an external system, let the agent check existing sources and schemas first.
  4. Follow the agent's guidance to pick the right source type and set up the import.

What this skill tells your AI

The instructions your AI receives, as published by posthog/posthog-foss in products/warehouse_sources/skills/suggesting-data-imports/SKILL.md and read by ahel’s review.

This skill helps identify when data the user needs lives outside PostHog and guides them toward importing it via the data warehouse. The key insight is recognizing the gap — then connecting it to the right source type.

What PostHog collects natively

PostHog collects product analytics events, persons, sessions, and groups via its SDKs. Additional products are available but must be enabled: session replay, feature flags, experiments, surveys, web analytics, error tracking, AI observability, conversations, logs, revenue analytics, workflows, CDP destinations, and batch exports. PostHog does not collect external business data like payments, subscriptions, CRM records, support tickets from other systems, or production database tables — that data must be imported via the data warehouse.

When to use this skill

  • A HogQL query fails because a table doesn't exist
  • The user asks about data from an external system (Stripe, Hubspot, Salesforce, etc.)
  • The user wants to correlate PostHog analytics with business data (revenue, support tickets, CRM records, etc)
  • The user asks "how do I get my X data into PostHog?"
  • Analysis requires joining PostHog events with external data
  • The user asks about exporting PostHog data for comparison elsewhere (in a google sheet, external warehouse, etc)

Workflow

1. Understand what data is missing

Listen for signals that the user needs external data:

  • They mention a specific tool or system (Stripe, Hubspot, Zendesk, their production database, etc.)
  • A query references a table that doesn't exist in PostHog
  • They want to analyze something PostHog doesn't track natively (revenue, support tickets, CRM deals, etc.)

If a query failed, check the error — if it's "table not found" or similar, the data likely needs to be imported.

2. Check what's already connected

Call posthog:external-data-sources-list to see existing sources. The data might already be imported but the user doesn't know the table name or prefix.

If a source exists for the system they're asking about, call posthog:external-data-schemas-list to show the available tables. The data might be there but under a different name or prefix.

Also query system.information_schema.tables with posthog:execute-sql to see all queryable tables — the data might already be available as a view or joined table.

3. Identify the right source type

If the data isn't imported yet, call posthog:external-data-sources-wizard to see available source types — when enumerating without source_type, pass fields: ['*.name', '*.caption'] to skip the large per-source config field definitions. Match the user's need to a source:

Common patterns:

User wantsSource typeKey tables
Revenue / payment dataStripe, Chargebee, Shopifycharges, subscriptions, invoices, customers
CRM / sales pipelineHubspot, Salesforce, Attiocontacts, deals, companies
Support ticketsZendesktickets, users, organizations
Product data from their DBPostgres, MySQL, BigQuery, Snowflake, Redshiftuser's own tables
Marketing / adsGoogle Ads, Meta Ads, LinkedIn Ads, TikTok Adscampaigns, ad_groups, ads
Email marketingMailchimp, Klaviyocampaigns, lists, subscribers
Project managementLinearissues, projects
Error tracking (external)Sentryissues, events

4. Suggest the import

Present the recommendation concisely:

  • What source type to connect
  • What tables would become available
  • How this enables the analysis they want

Example: "Your Stripe data isn't in PostHog yet. If you connect a Stripe source, you'll get tables like charges, subscriptions, and customers that you can join with PostHog events to analyze revenue by user behavior."

5. Offer to set up the source

If the user wants to proceed, the fastest path is the one-step data-warehouse-source-setup tool (validate creds → discover tables → sync defaults → create, in one call), with data-warehouse-source-connect-link to collect credentials securely in the browser rather than in chat. For anything beyond the happy path (hand-picking tables, non-default sync types, webhooks, CDC), hand off to the setting-up-a-data-warehouse-source skill, which covers the full flow, sync-type selection, webhook registration, and prefix guidance. Do not duplicate that workflow here.

6. Show what's possible after import

Once connected, help the user write their first query joining PostHog data with the imported data. Use posthog:execute-sql to demonstrate.

Common join patterns:

  • Join Stripe customers with PostHog persons on email: SELECT * FROM stripe_customers sc JOIN persons p ON sc.email = p.properties.$email
  • Join CRM deals with events: correlate product usage with sales outcomes
  • Join support tickets with session recordings: find recordings for users who filed tickets

Important notes

  • Don't guess table names. Always check system.information_schema.tables (via posthog:execute-sql) and posthog:external-data-schemas-list before saying data doesn't exist.
  • Check prefixes. Imported tables are often prefixed (e.g. stripe_charges not charges). The user might not know the prefix.
  • Collect credentials securely. Use data-warehouse-source-connect-link to hand the user a browser link — it opens a minimal connect page rendering the source's full connection form (OAuth or credentials, whichever the source offers) that stashes the details temporarily without creating the source. Afterwards pass {"credential_id": <id>} (discovered via data-warehouse-stored-credentials-list) to data-warehouse-source-setup — stored credentials are single-use and expire after 24 hours. Don't collect passwords or OAuth tokens in chat.
  • Not all systems are supported. If the user's system isn't in the wizard list, suggest using Postgres/MySQL as a bridge if they can export to a database, or mention that custom sources can be requested.
  • Connecting a source also documents it. After the first sync, PostHog automatically generates semantic descriptions for the imported tables and columns (from the source database's own column comments where present, plus an LLM pass using the table relationships and the team's business context). Those descriptions surface in system.information_schema.columns (query it with posthog:execute-sql), so once a source is connected the agent can reason about what each column means and how tables join — not just their names and types. Mention this when recommending an import: connecting the source is what makes the data answerable.

Related tools

  • posthog:external-data-sources-list: Check existing source connections
  • posthog:external-data-schemas-list: Check what tables are already imported
  • posthog:execute-sql over system.information_schema.*: See all queryable tables including views
  • posthog:external-data-sources-wizard: Get available source types (pass fields: ['*.name', '*.caption'] when enumerating)
  • posthog:data-warehouse-source-connect-link: Get a secure browser/OAuth link to collect credentials
  • posthog:data-warehouse-source-setup: One-step create (validate, discover tables, apply sync defaults, create)
  • posthog:execute-sql: Run queries to demonstrate what's possible

Related skills

  • setting-up-a-data-warehouse-source: Full source creation workflow — hand off here once the user decides to connect a source

Signals

GitHub stars
721
Forks
120
Last commit
Oct 2026

Others that do the same job

Questions

What data does PostHog collect natively?
PostHog collects product analytics events, persons, sessions, and groups via its SDKs. Other products like session replay, feature flags, and surveys must be enabled. External business data such as payments, CRM records, and production database tables is not collected natively.
When should this skill be used?
Use it when a HogQL query fails because a table does not exist, when the user asks about data from an external system, when they want to correlate PostHog analytics with business data, or when they ask how to get external data into PostHog.
Advanced
Item type
skill
Key
suggesting-data-imports
Source
github.com/posthog/posthog-foss