Authoring Airflow DAGs
SkillProductivityWrite production-grade Apache Airflow DAGs using the TaskFlow API, idempotent tasks, correct scheduling and catchup, retries/SLAs, connections/variables, and avoiding top-level code. Use when creating or reviewing Airflow DAGs, scheduling pipelines, wiring task dependencies, configuring retries/backfills, or fixing non-idempotent tasks.
Use Authoring Airflow DAGs in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Authoring Airflow DAGs and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Authoring Airflow DAGs skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by unknown-333/awesome-data-engineering-skills in skills/authoring-airflow-dags/SKILL.md and read by Ahel’s review.
When to use
- Creating or refactoring Airflow DAGs and tasks.
- Configuring schedules, catchup/backfill, retries, and SLAs.
- Passing data between tasks (XCom) or using connections/variables.
- Do NOT use for diagnosing a broken running DAG (use
debugging-airflow-pipelines).
Workflow
- [ ] Make each task idempotent and parameterized by the data interval
- [ ] Keep expensive/import-heavy code inside tasks, not at module top level
- [ ] Set schedule + catchup deliberately
- [ ] Configure retries, retry_delay, and SLAs
- [ ] Wire dependencies via TaskFlow return values or >> operators
- Idempotent tasks — a task for the
2026-01-15interval must produce the same result whether it runs once or is re-run. Use the data interval, notdatetime.now(). - No heavy top-level code — the scheduler parses every DAG file frequently; database calls, API calls, or big imports at module level slow scheduling and can break parsing. Put them inside tasks.
- Schedule + catchup on purpose —
catchup=Truebackfills every missed interval fromstart_date; default toFalseunless you want that. - Retries and SLAs — transient failures are normal; set
retriesandretry_delay; use SLAs/alerts for lateness.
Patterns
TaskFlow DAG, idempotent and cleanly wired:
from airflow.decorators import dag, task
import pendulum
@dag(
schedule="@daily",
start_date=pendulum.datetime(2026, 1, 1, tz="UTC"),
catchup=False,
default_args={"retries": 3, "retry_delay": pendulum.duration(minutes=5)},
tags=["orders"],
)
def orders_pipeline():
@task
def extract(data_interval_start=None):
# Use the interval, not now(), so re-runs are deterministic.
return fetch_orders(day=data_interval_start.date())
@task
def load(rows):
# Delete-insert the partition -> idempotent on retry.
overwrite_partition("fct_orders", rows)
load(extract())
orders_pipeline()
Pass small data via XCom (return values); pass large data via storage — write to S3/GCS/warehouse and pass the path/key, never megabytes through XCom.
Use connections/variables for secrets and config (BaseHook.get_connection,
Variable.get), never hard-coded credentials.
Common pitfalls
- Top-level API/DB calls or heavy imports — slow the scheduler and can fail DAG parsing across the whole deployment.
datetime.now()inside tasks — breaks idempotency and backfills; usedata_interval_start/_end.catchup=Trueunintentionally — floods the cluster with historical runs on first deploy.- Large payloads through XCom — bloats the metadata DB; pass references.
- Dynamic
start_date(e.g.days_ago) — makes schedules nondeterministic; use a fixed timestamp. - One monster task — split extract/transform/load so retries are granular.
References
Signals
- GitHub stars
- 21
- Last commit
- Aug 2026
Advanced
- Item type
- skill
- Key
authoring-airflow-dags- Source
- github.com/unknown-333/awesome-data-engineering-skills
github.com/unknown-333/awesome-data-engineering-skills
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonsocial
Skill · coreyhaines31
More in Productivitygws-calendar
Skill · googleworkspace
More in Productivitylark-minutes
Skill · larksuite
More in Productivitylark-workflow-standup-report
Skill · larksuite
More in Productivity