- What Is a Python ETL Framework? {#what-is}
- Why Use Python for ETL Pipelines in 2026? {#why}
- What to Look for in a Python ETL Framework {#what-to-look-for}
- Which Are the Best Python ETL Frameworks? {#best}
- Comparison of Python ETL Frameworks {#comparison}
- Other ETL Tools, Libraries & Frameworks {#other}
- When to Complement Python ETL with a Low-Code Platform {#integrate-io}
- Frequently Asked Questions
A Python ETL framework is a library, package, or orchestration tool that uses Python to extract data from source systems, apply transformations, and load it into a destination such as a data warehouse or database. Unlike visual drag-and-drop platforms, Python ETL frameworks give data engineers direct control over pipeline logic, scheduling, and transformation rules through code. Most production pipelines combine an orchestration framework (like Airflow or Dagster) with a transformation library (like pandas or petl).
Quick Answer: Best Python ETL Frameworks in 2026
- Apache Airflow, best for orchestrating complex, multi-step pipelines
- Dagster, best for data lineage, asset-based workflows, and testability
- Prefect, best for teams migrating from Airflow who want a managed option
- Luigi, best for lightweight, dependency-based batch workflows
- pandas, best for in-memory data wrangling on small to medium datasets
- petl, best for simple, scriptable data cleanup tasks
- Bonobo, best for quick-start ETL prototyping with minimal setup
- Not a Python coder? Integrate.io delivers ETL, ELT, CDC, and Reverse ETL with 150+ connectors and no code required.
Key Takeaways
- Python ETL frameworks split into two categories: orchestration tools (Airflow, Dagster, Prefect, Luigi) and transformation libraries (pandas, petl, Bonobo). Most teams use both.
- Dagster and Prefect are actively displacing Airflow in new builds due to better local development experience and built-in observability.
- pandas is the most widely used Python library for data manipulation, but it works in-memory and is not suitable for large-scale datasets.
- None of these tools offer built-in connectors, Change Data Capture (CDC), or Reverse ETL out of the box. For that, pair them with a platform like Integrate.io.
- Bubbles has been removed from the main list; it is no longer actively maintained as of 2026.
- On average, it takes 8 weeks to learn the basics of Python. If your team lacks that bandwidth, a low-code platform is worth evaluating.
What Is a Python ETL Framework? {#what-is}
Extract, transform, load (ETL) is a critical component of data engineering, enabling efficient data transfer between systems. A Python ETL framework provides the scaffolding to automate that process using Python code.
These frameworks fall into two categories:
- Orchestration frameworks (Airflow, Dagster, Prefect, Luigi): Handle scheduling, dependency management, retries, and monitoring across pipeline steps.
- Transformation libraries (pandas, petl, Bonobo): Handle the actual manipulation of data rows, columns, and formats.
Most production data pipelines use both: an orchestration layer to coordinate when and how jobs run, and a transformation library to process the data itself.
When Python ETL makes sense:
- Your team has Python experience and needs custom pipeline logic.
- You have unique transformation requirements that prebuilt connectors cannot handle.
- You want full control over scheduling, retries, and execution order.
Why Use Python for ETL Pipelines in 2026? {#why}
Python is the dominant language for ETL pipelines in 2026. It handles complex schemas and large volumes of big data well, has an enormous active community, and offers more flexibility than most visual ETL tools. You can build a pipeline tailored exactly to your needs, from simple CSV transformations to distributed processing across cloud warehouses.
That said, Python ETL is not always the right choice. It requires coding skills, ongoing maintenance, and careful dependency management. For teams without dedicated data engineering resources, a low-code platform often delivers faster results.
Python ETL makes the most sense when:
- You have experience with Python and want to build pipelines from scratch.
- You have simple ETL requirements and want a lightweight solution.
- You have a unique need that can only be met by custom-coded pipeline logic.
There are more than a hundred Python tools available for ETL in 2026, including frameworks, libraries, and full orchestration platforms. The eight covered below were selected based on usability, active maintenance, and relevance to modern data transformation workflows.
Recommended Reading: Building an ETL Pipeline in Python
What to Look for in a Python ETL Framework {#what-to-look-for}
Not every Python ETL framework fits every team. Before committing, evaluate these six criteria.
Orchestration vs. Transformation Focus
Orchestration tools (Airflow, Dagster, Prefect) manage scheduling, dependencies, and retries. Transformation libraries (pandas, petl) manipulate data. Most teams need both. Identify your primary gap before choosing a tool.
Scale Requirements
Libraries like pandas and petl work in-memory and are not suitable for datasets that exceed available RAM. For large-scale distributed processing, PySpark or a managed cloud service is a better fit. Understand your data volumes before committing to an in-memory tool.
Ease of Setup
Self-hosted orchestration tools like Airflow require infrastructure setup, worker configuration, and ongoing maintenance. Managed options (Astronomer, Prefect Cloud) reduce that overhead significantly. Factor in your team's DevOps capacity.
Real-Time vs. Batch
Most Python ETL frameworks are batch-only. If you need sub-minute data replication or streaming, Python frameworks are generally not the right fit without significant custom engineering. Purpose-built CDC platforms handle this more reliably.
Community and Maintenance Status
An unmaintained framework is a liability. Check the GitHub commit history and last release date before adopting any tool. This is why Bubbles has been removed from the main list in this guide; it is no longer actively maintained as of 2026.
Cloud Warehouse Compatibility
If your destination is Snowflake, BigQuery, or Redshift, confirm the framework has native connectors or well-maintained community plugins. Airflow and Dagster both have strong cloud warehouse support. Libraries like pandas connect via SQLAlchemy or native drivers.
When to Pair with a Low-Code Platform
Python frameworks handle pipeline logic well but require code for every connector, transformation, and schedule. If your team needs 150+ prebuilt connectors, built-in data orchestration, or Reverse ETL without writing pipeline logic from scratch, a platform like Integrate.io complements Python workflows effectively. More on this in the closing section.
Which Are the Best Python ETL Frameworks? {#best}
Apache Airflow, Dagster, and Prefect are among the best Python ETL frameworks for orchestrating complex data pipelines. For transformation tasks, pandas and petl remain the most widely used options. For teams preferring a hybrid model, Integrate.io complements Python-based workflows by allowing external Python scripts to run within its low-code pipeline architecture.
1. Apache Airflow for Python-Based Workflows {#airflow}
Best for: Orchestrating complex, multi-step pipelines with scheduling, monitoring, and cloud integrations.
Apache Airflow is an open-source, Python-based workflow automation tool used for setting up and maintaining powerful data pipelines. It is not an ETL tool in the traditional sense, but it manages, structures, and organizes ETL pipelines using Directed Acyclic Graphs (DAGs). DAGs form relationships and dependencies between tasks, allowing a single branch to run multiple times or skip branches when necessary.
For example, you can make task A run after task B, and make task C run every 2 minutes. Or make task A and task B run every 2 minutes and make task C run after task B.
The typical Airflow architecture:
Metadata database > scheduler > executor > workers
The metadata database stores workflows and tasks (DAGs). The scheduler uses DAG definitions to select tasks. The executor determines which worker runs each task. Workers execute the logic of workflows and tasks.
Airflow has hooks and operators for Google Cloud and AWS, making it well-suited for cloud warehousing environments. It is not a library, so it needs to be deployed, which may not be practical for small ETL jobs.
Airflow makes the most sense for long ETL jobs or pipelines with multiple dependent steps.
Facts about Apache Airflow:
- It won InfoWorld's 2020 Best of Open Source Software Award.
- The majority of Airflow users leverage Celery to simplify execution management.
- You can schedule automated DAG workflows via the Airflow WebUI.
- Airflow uses a command-line interface, which is useful for isolating execution tasks outside of scheduler workflows.
G2 rating: 4.4/5
Advantages:
- Python-based DAGs enable full customization
- Strong integration support across cloud and data platforms
- Built-in scheduler, monitoring, and web UI
Limitations:
- Clunky and unintuitive UI
- Requires Python coding skills
- Not designed for real-time streaming
- Can become complex to scale with large DAGs
Pricing:
- Free if self-hosted
- Managed versions (Astronomer, Google Composer, AWS MWAA) charge based on usage or per-hour compute
2. Dagster for Asset-Based Data Orchestration {#dagster}
Best for: Teams that need data lineage, asset-based pipelines, and testability in production environments.
Dagster is a modern open-source data orchestration platform that has gained significant adoption as teams look for Airflow alternatives with better observability and local development experience. Where Airflow organizes work around tasks, Dagster organizes work around software-defined assets: the data outputs your pipelines produce. This makes it easier to track data lineage, understand dependencies, and test pipelines before deploying them.
Dagster's asset-centric model is a meaningful shift from DAG-centric tools. Instead of asking "what tasks ran?", you ask "what data was produced, and is it fresh?" This approach maps more naturally to how analytics and data science teams think about their work.
Key differentiators vs. Airflow:
- Asset-centric model with built-in data lineage tracking
- Better local development and testing experience
- Built-in observability without additional tooling
- Dagster Cloud offers a managed option with a free tier
Facts about Dagster:
- Dagster is actively maintained with frequent releases and a growing community.
- It integrates natively with Snowflake, BigQuery, dbt, and Spark.
- Dagster's Launchpad UI provides a cleaner development experience than Airflow's web UI.
- It supports partitioned assets, making incremental processing straightforward.
G2 rating: 4.4/5
Advantages:
- Asset-centric pipeline model improves lineage and observability
- Excellent local development and testing workflow
- Strong integrations with modern data stack tools
- Managed cloud option available
Limitations:
- Steeper learning curve than simpler tools like Luigi or Bonobo
- Smaller community than Airflow
- Asset model requires a mindset shift for teams coming from task-based orchestrators
Pricing:
- Free and open source (self-hosted)
- Dagster Cloud has a free tier; paid plans based on usage
3. Prefect for Modern Workflow Orchestration {#prefect}
Best for: Teams migrating from Airflow who want simpler DAG management and a hosted option.
Prefect is a modern Python-based workflow orchestration platform designed to address Airflow's pain points. It uses a Python-native API that feels more intuitive than Airflow's DAG syntax, and it supports both local execution and a fully managed cloud option (Prefect Cloud). Prefect is often the easiest migration path for Airflow users who want to reduce infrastructure overhead without rewriting all their pipeline logic.
Prefect's key innovation is its "negative engineering" philosophy: it handles failures, retries, and observability so you can focus on the logic of your pipelines rather than the plumbing.
Key differentiators vs. Airflow:
- Simpler Python-native API, no DAG file structure required
- Prefect Cloud provides a managed option with a free tier
- Better handling of dynamic, parametrized workflows
- Faster local testing without a full Airflow deployment
Facts about Prefect:
- Prefect supports both Prefect Cloud (managed) and Prefect Server (self-hosted).
- It integrates with Snowflake, BigQuery, AWS, and GCP natively.
- Prefect flows are plain Python functions decorated with @flow, making onboarding straightforward for Python developers.
- It has strong support for event-driven and scheduled workflows.
G2 rating: 4.3/5
Advantages:
- Python-native API is easier to learn than Airflow DAGs
- Managed cloud option reduces infrastructure burden
- Strong observability and alerting built in
- Active development and growing community
Limitations:
- Prefect Cloud costs can scale with usage
- Smaller ecosystem of community integrations than Airflow
- Self-hosted Prefect Server requires infrastructure management
Pricing:
- Prefect Server: free and open source
- Prefect Cloud: free tier available; paid plans based on usage
4. Luigi for Complex Python Pipelines {#luigi}
Best for: Simple, dependency-based batch workflows where lightweight setup is a priority.
Luigi is an open-source, Python-based tool for building complex pipelines. Developed by Spotify to automate heavy workloads, Luigi is used by data-driven organizations such as Stripe and Red Hat.
Three main benefits of Luigi:
- Dependency management with visualization
- Failure recovery through checkpoints
- Command-line interface integration
The primary difference between Luigi and Airflow is how they execute tasks and dependencies. In Luigi, you work with "tasks" and "targets," and tasks consume targets. This target-based approach works well for simple Python-based ETL, but Luigi can struggle with highly complex tasks.
Luigi is best suited for automating simple ETL processes such as log processing. Its strict pipeline-like structure limits its ability to handle complex tasks, and even simple processes require Python coding skills.
Facts about Luigi:
- Once Luigi runs tasks, you cannot interact with the processes.
- Unlike Airflow, Luigi does not automatically schedule, alert, monitor, or sync tasks to workers.
- Towards Data Science describes Luigi's approach as "target-based" with a "minimal" UI and no user interaction for running processes.
G2 rating: N/A
Advantages:
- Simple task and dependency model
- Lightweight UI for basic visualization
- Reliable with checkpointing and recovery features
Limitations:
- Very basic UI
- Not ideal for large-scale workflows
- Less extensible than modern alternatives like Dagster or Prefect
Pricing:
- Free and open source
5. pandas for Data Structures and Analysis {#pandas}
Best for: In-memory data wrangling and transformation on small to medium datasets.
pandas is a widely used open-source library that provides data structures and data analysis tools for Python. It is particularly useful for ETL tasks, as it adds R-style data frames that make processes like data cleansing and transformation easier. With pandas, you can load data from various sources, clean and transform it, and write it to Excel, CSV, or SQL databases with minimal code.
pandas is not the best choice for large-scale data processing. It works entirely in-memory, so datasets that exceed available RAM will cause performance issues or failures. For large-scale processing, tools like PySpark or a managed cloud service are better fits.
When pandas makes sense:
pandas is ideal for ETL tasks involving small to medium datasets where the primary focus is data cleansing, transformation, and manipulation. It is particularly useful when extracting data, cleaning it, and writing it to Excel, CSV, or an SQL database.
Facts about pandas:
- NumFocus sponsors pandas.
- Many Python users choose pandas for ETL batch processing.
- pandas holds a 4.6/5 rating on G2. Users call it "powerful" and "very practical," though note there is "a learning curve."
G2 rating: 4.6/5
Advantages:
- Powerful DataFrame API for data manipulation and analysis
- Large community and excellent documentation
- Rich set of built-in functions
Limitations:
- Works in-memory, not suitable for large datasets
- High memory consumption
- Performance drops as data size increases
Pricing:
- Free and open source
6. petl as a Python ETL Solution {#petl}
Best for: Basic ETL functionality where simplicity and low memory use matter more than speed.
petl is among the most straightforward Python ETL frameworks available. It is a widely used open-source library that simplifies building tables, extracting data from various sources, and performing standard ETL tasks. It is similar in functionality to pandas but without the same level of data analysis capabilities.
petl handles complex datasets efficiently by utilizing system memory carefully and providing reliable scalability for its scope. It is not as fast as other ETL tools, but it is easier to use than building ETL with SQLAlchemy or other custom-coded solutions.
When petl makes sense:
petl is a good choice when you need ETL basics without advanced analytics, and speed is not a critical factor.
Facts about petl:
- petl stands for "Python ETL."
- It is not known for speed or handling of large datasets.
- Towards Data Science praises petl for its support of standard transformations including sorting, joining, aggregation, and row operations.
G2 rating: N/A
Advantages:
- Very lightweight with low memory use
- Simple API for extract, transform, load operations
- Supports a wide range of data formats
Limitations:
- Limited transformation capabilities
- No orchestration or scheduling support
- Not suitable for complex pipelines
Pricing:
- Free and open source
7. Bonobo as a Lightweight Python ETL Framework {#bonobo}
Best for: Quick-start ETL prototyping with minimal API learning overhead.
Bonobo is a lightweight Python ETL framework that allows for rapid deployment of data pipelines and parallel execution. It supports CSV, JSON, XML, XLS, and SQL, and adheres to atomic UNIX principles. One of its main benefits is that it requires minimal learning of a new API, making it accessible to those with basic Python knowledge. It can build graphs, create libraries, and automate simple ETL batch processes.
When Bonobo makes sense:
Bonobo is ideal for simple, lightweight ETL jobs where you do not have time or resources to learn a new API. It still requires basic Python knowledge. For teams looking for a no-code solution, Integrate.io is a more practical option.
Facts about Bonobo:
- Bonobo offers a Docker extension for running jobs within Docker containers.
- It has a command-line interface (CLI) for execution.
- Bonobo has built-in Graphviz support for visualizing ETL job graphs.
- It also has an SQLAlchemy extension (currently in alpha).
- A writeup describes writing a first ETL job in Bonobo as "simple and straightforward."
G2 rating: N/A
Advantages:
- Clean and easy-to-write ETL pipelines
- Modular and reusable components
- Supports parallel execution
Limitations:
- Lacks built-in scheduling
- Limited ecosystem and integrations
- Not scalable for large workflows
Pricing:
- Free and open source
Comparison of Python ETL Frameworks {#comparison}
| Feature/Aspect | Apache Airflow | Dagster | Prefect | Luigi | pandas | petl | Bonobo |
|---|---|---|---|---|---|---|---|
| Type | Workflow orchestration | Asset-based orchestration | Workflow orchestration | Task orchestration | Data transformation library | Lightweight ETL library | Lightweight ETL framework |
| Best For | Complex multi-step pipelines | Data lineage and testability | Airflow migration, managed option | Simple dependency-based batch | In-memory data wrangling | Basic scriptable ETL | Quick-start ETL prototyping |
| Ease of Use | Moderate, requires Python | Moderate, asset model learning curve | Easy Python-native API | Simple API, requires code | Easy for Python users | Very easy, procedural logic | Easy, minimal API |
| Transformation Support | Indirect via Python operators | Yes, via asset computations | Yes, via flow tasks | Yes, custom task logic | Extensive DataFrame transforms | Basic row-wise transforms | Yes, via graph nodes |
| Real-Time Capabilities | No | No | No | No | No | No | No |
| Scheduling | Yes, built-in | Yes, built-in | Yes, built-in | Yes, internal scheduler | No | No | No |
| Parallelism | Yes (Celery, Kubernetes) | Yes | Yes | Yes | Limited | No | Yes, multi-threading |
| Cloud Warehouse Support | Strong (native connectors) | Strong (native connectors) | Strong (native connectors) | Limited (custom plugins) | Via SQLAlchemy/drivers | Via SQLAlchemy/drivers | Limited |
| Active Maintenance | Yes | Yes | Yes | Yes | Yes | Yes | Limited |
| Pricing | Free / managed options | Free / cloud tier | Free / cloud tier | Free | Free | Free | Free |
| Support | Large community, commercial options | Active community, growing | Active community, commercial | Community support | Massive community | Small community | Niche community |
Other ETL Tools, Libraries & Frameworks {#other}
There are too many Python ETL tools to cover in a single list. Below are additional options worth knowing, organized by category.
Python
- BeautifulSoup: Pulls data out of webpages (XML, HTML) and integrates with tools like petl.
- PyQuery: Extracts data from webpages with a jQuery-like syntax.
- Blaze: An interface for querying data, part of the Blaze Ecosystem alongside Dask, Datashape, DyND, and Odo.
- Dask: Parallel computing via task scheduling; also processes continuous data streams.
- Datashape: A data-description language for in-situ structured data.
- DyND: Python exposure for DyND, a C++ library for dynamic multidimensional arrays.
- Odo: Moves data between multiple containers using native CSV loading capabilities.
- Joblib: Python functions for pipelines with strong support for NumPy arrays.
- lxml: Processes HTML and XML in Python.
- Retrying: Adds retry behavior to Python executions.
- riko: A Yahoo! Pipes replacement useful for stream data extraction.
- Bubbles (bubbles.databrewery.org): Metadata-driven ETL framework; no longer actively maintained as of 2026. Open Knowledge Labs described it as "a framework for ETL written in Python, but not necessarily meant to be used from Python only." Use a maintained alternative for new projects.
Cloud-Based
- Integrate.io: Point-and-click, 200+ out-of-the-box integrations, Salesforce-to-Salesforce integration, and more. No code required. See the closing section for details on Integrate.io's ETL platform capabilities.
- AWS Data Pipeline: Amazon's data pipeline solution for AWS instances and legacy servers.
- AWS Glue: Amazon's fully managed ETL solution, managed through the AWS Management Console.
- AWS Batch: Batch computing on AWS resources with good scalability for large jobs.
- Google Dataflow: Google's ETL solution for batch jobs and streams.
- Azure Data Factory: Microsoft's ETL solution.
Miscellaneous
- Toil: Handles ETL almost identically to Luigi. Use Luigi to wrap Toil pipelines for additional checkpointing.
- Pachyderm: Another Airflow alternative. A writeup covers the key differences between Airflow and Pachyderm.
- Mara: Sits between pure Python and Apache Airflow; fast and simple to set up.
- Pinball: Pinterest's workflow manager with auto-retries, priorities, and horizontal scalability.
- Azkaban: Created by LinkedIn, a Java-based tool for Hadoop batches. For hyper-complex Hadoop batches, consider Oozie instead.
- Dray.it: A Docker workflow engine for resource management.
- Spark: Full batch streaming ETL at scale.
ETL tools also exist in other languages. Java options include Spring Batch, Scriptella, Apatar, Jaspersoft ETL, EasyBatch, GETL, and JSR 352. Go has Crunch. JavaScript (Node.js) options include NoFlo, Extraload, Empujar, Datapumps, and proc-that. Python remains the dominant language for modern data pipeline development.
When to Complement Python ETL with a Low-Code Platform {#integrate-io}
Python ETL frameworks give you control and flexibility. But they require coding skills, ongoing maintenance, and custom connector development for every new data source. For many teams, that overhead is the bottleneck, not the pipeline logic itself.
Integrate.io's low-code data pipeline platform is built for exactly this scenario. It complements Python-based workflows by handling the connectors, scheduling, and infrastructure so your team can focus on the logic that actually requires custom code.
When Integrate.io makes sense alongside Python ETL:
- Your team needs 150+ prebuilt connectors without writing custom integration code.
- You need Change Data Capture (CDC) with sub-60-second replication to power real-time dashboards.
- You need Reverse ETL to push warehouse data back into operational tools like Salesforce.
- Non-technical team members need to build or monitor pipelines without Python expertise.
- You want data orchestration with built-in alerting, scheduling, and monitoring without deploying Airflow infrastructure.
Integrate.io also supports running external Python scripts within its pipeline architecture, so teams do not have to choose between low-code convenience and custom logic. You get both.
It is SOC 2 certified, GDPR, HIPAA, and CCPA compliant, and backed by 24/7 support from a dedicated team. Customers include Samsung, Caterpillar, and 7-Eleven.
Integrate.io's Integrate.io's ETL features include 220+ built-in transformations, a visual pipeline builder, and a fixed-fee unlimited usage plan starting at $1,999/month with no row limits and no pipeline caps.
Schedule a demo with our team to see if Integrate.io is the right fit for your data stack.
Frequently Asked Questions
What is a Python ETL framework?
A Python ETL framework is a library, package, or orchestration tool that uses Python to extract data from source systems, transform it, and load it into a destination such as a data warehouse or database. Frameworks like Airflow and Dagster handle orchestration and scheduling; libraries like pandas and petl handle data manipulation. Most production pipelines use both.
What is the difference between a Python ETL framework and a Python ETL library?
Frameworks (Airflow, Dagster, Prefect) handle orchestration and scheduling: they manage when tasks run, handle retries, and track dependencies. Libraries (pandas, petl) handle data manipulation: they transform rows, columns, and formats. Most production pipelines use both a framework for coordination and a library for transformation logic.
Can I use Python ETL tools with a cloud data warehouse like Snowflake or BigQuery?
Yes. Tools like Airflow and Dagster have native connectors for Snowflake, BigQuery, and Redshift. Libraries like pandas connect via SQLAlchemy or native database drivers. Prefect also has strong cloud warehouse integrations out of the box.
When should I use a low-code platform instead of a Python ETL framework?
When your team lacks Python expertise, needs faster deployment, or requires 150+ prebuilt connectors without writing custom code. Platforms like Integrate.io handle ETL, ELT, Change Data Capture (CDC), and Reverse ETL without writing pipeline logic from scratch. They also work well alongside Python frameworks for teams that need both.
What are the top Python ETL tools suitable for non-technical users?
Most Python ETL frameworks require coding skills. For non-technical users, the best options are managed platforms with low-code interfaces. Integrate.io is purpose-built for this: it offers a visual pipeline builder, 220+ transformations, and 150+ connectors with no Python required. Among Python-native tools, Bonobo has the simplest syntax, but it still requires basic Python knowledge.
What are the most popular Python libraries for building ETL pipelines?
The most commonly used libraries include pandas for data manipulation, SQLAlchemy for database connectivity, Airflow or Prefect for orchestration, and Dask for distributed processing. These tools offer flexibility for building custom, scalable ETL pipelines tailored to specific use cases.
Is Apache Airflow still the best Python orchestration tool in 2026?
Airflow remains the most widely deployed Python orchestration tool, but Dagster and Prefect are actively displacing it in new builds. Dagster's asset-centric model and better local development experience appeal to teams starting fresh. Prefect is the most common migration path for teams moving off Airflow. For existing Airflow deployments, the switching cost is real; for new projects, evaluate all three.
Why was Bubbles removed from the main list?
Bubbles is no longer actively maintained as of 2026. Including a deprecated tool in a "top frameworks" list would give misleading guidance. It has been moved to the "Other Tools" section with a note. If you are evaluating Python ETL frameworks for production use, choose a tool with active development and community support.