r/learnpython • • 22h ago

looking for optimizing my database bootstrapper which is taking >30 seconds.

Hello fellow devs i am building a FastAPI + SQLAlchemy 2.0 (async) with a metadata driven system which has lost of dynamic schema stuff based on the data tables

Right now my custom bootstrapper handles botjh creating the DDL Schema and a massive amount of Data seeding ( lots of data) sequentially. because it's doing manual idempotency checks (SELECT before INSERT) and running everything one by one, the whole provisioning process takes more than 30 seconds from scratch.

I'm looking into splitting the data fixture side out to a library like sqlalchemy-seedling to run independent seeders in parallel using asyncio.gather. But before I start rewriting, how do you all handle heavy initial database provisioning in async Python?

0 Upvotes

5 comments sorted by

2

u/neums08 20h ago

Use batch inserts with ON CONFLICT clauses. Run the seeding process in a fastapi BackgroundTask.

1

u/OkTill2666 19h ago

yep batching alone would be a huge win here before even thinking about parallelism

1

u/a18618 4h ago

Before rewriting around asyncio.gather, I'd measure where the 30 seconds actually go — DDL versus inserts — because parallelism only helps if the bottleneck is client-side. Against a single database, N concurrent seeders usually just queue on the same locks and the wall clock barely moves; batching, one transaction, and building indexes after the load are the usual big wins. Parallel seeders are a fine architecture, but they're an answer to a throughput problem you haven't confirmed you have yet.

0

u/ectomancer 20h ago

lazy import in Python 3.15