seedgraph¶
Seed your SQLAlchemy models as a referentially-consistent graph — one call, shared parents, verified links.
seedgraph fills your test database with a coherent graph of objects, in a single call. You declare what you want — "3 users, each with 2 posts, each post with 3 comments" — and the library builds the objects, links them, writes them to your session and then verifies every foreign key against the row it points at. Values are realistic and reproducible ("Jose Bishop", not "user-0"), generated by Faker with a fixed seed, valid for each column's type, and unique where the schema says so — even against rows already in the database. You can pin any column, replace how any column is generated, and attach the new graph to rows you already have. It plugs into pytest with no configuration.
Status: alpha — the API may still change before 1.0. Every guarantee below is backed by a named test, on SQLite and on PostgreSQL.
Documentation: jrachid.github.io/seedgraph — recipes for pytest, FastAPI and custom column types, the API reference, and a Claude Code plugin that teaches coding agents to use seedgraph.
Quick start¶
pip install seedgraph
pip install "seedgraph[async]" # for seed_async and the async pytest fixtures
Requires Python 3.11+, SQLAlchemy 2.x and Faker 30+. The async extra adds greenlet (through sqlalchemy[asyncio]) and aiosqlite.
Declare a shape from a root model; each key walks a one-to-many or many-to-many relationship, by relationship name or by target class name:
from sqlalchemy import ForeignKey, create_engine
from sqlalchemy.orm import DeclarativeBase, Mapped, Session, mapped_column, relationship
from seedgraph import seed
class Base(DeclarativeBase):
pass
class User(Base):
__tablename__ = "users"
id: Mapped[int] = mapped_column(primary_key=True)
name: Mapped[str]
email: Mapped[str] = mapped_column(unique=True)
posts: Mapped[list["Post"]] = relationship(back_populates="author")
class Post(Base):
__tablename__ = "posts"
id: Mapped[int] = mapped_column(primary_key=True)
title: Mapped[str]
subtitle: Mapped[str | None]
author_id: Mapped[int] = mapped_column(ForeignKey("users.id"))
author: Mapped[User] = relationship(back_populates="posts")
comments: Mapped[list["Comment"]] = relationship(back_populates="post")
class Comment(Base):
__tablename__ = "comments"
id: Mapped[int] = mapped_column(primary_key=True)
body: Mapped[str]
post_id: Mapped[int] = mapped_column(ForeignKey("posts.id"))
post: Mapped[Post] = relationship(back_populates="comments")
engine = create_engine("sqlite://")
Base.metadata.create_all(engine)
session = Session(engine)
graph = seed(session, User, post=2, post__comment=3) # 3 users by default
assert len(graph.users) == 3
assert len(graph.posts) == 6
assert len(graph.comments) == 18
post = graph.users[0].posts[0]
assert post.author_id == graph.users[0].id # real key, assigned by the database
seed() returns once the graph is flushed and verified; commit or roll back as your test needs. seed_async(async_session, ...) is its twin for an AsyncSession. The graph exposes every table of the model's metadata as an attribute, empty when the shape built none.
Existing and missing parents¶
alice = session.get(User, 1)
graph = seed(session, Post, parents=[alice]) # every post's author is alice; alice is not in graph.users
graph = seed(session, Comment) # one Post and one User are generated, shared by all comments
A link first takes the nearest ancestor of its type in the shape, then the object of that type passed in parents (optional links included). A required link still empty gets one generated parent per type, shared by every object that needs it. Several objects of one type are accepted in parents; a link towards a single parent refuses to choose between them with AmbiguousParentError.
Many-to-many¶
from sqlalchemy import Column, Table
article_tag = Table(
"article_tag",
Base.metadata,
Column("article_id", ForeignKey("articles.id"), primary_key=True),
Column("tag_id", ForeignKey("tags.id"), primary_key=True),
)
class Tag(Base):
__tablename__ = "tags"
id: Mapped[int] = mapped_column(primary_key=True)
name: Mapped[str] = mapped_column(unique=True)
class Article(Base):
__tablename__ = "articles"
id: Mapped[int] = mapped_column(primary_key=True)
title: Mapped[str]
tags: Mapped[list[Tag]] = relationship(secondary=article_tag)
Base.metadata.create_all(engine)
python, sql = Tag(name="python"), Tag(name="sql")
session.add_all([python, sql])
graph = seed(session, Article, article=5, tags=3) # 15 new tags, 3 per article
assert len(graph.tags) == 15
graph = seed(session, Article, article=5, parents=[python, sql]) # every article tagged with both existing tags
assert all(article.tags == [python, sql] for article in graph.articles)
A count keeps its one-to-many meaning: new objects for each parent. Objects passed in parents join every many-to-many collection of their type, next to the ones the shape builds. SQLAlchemy writes the association rows itself.
Pinning and generating values¶
graph = seed(
session,
User,
post=2,
generators={User: {"name": lambda ctx: ctx.fake.first_name()}}, # replace how a column is generated
overrides={Post: {"title": "Imposed", "subtitle": None}}, # pin a value, None included
)
ctx.fake is the session's seeded Faker; ctx.column is the column name. An override can also be a callable taking the same context.
pytest¶
Installing seedgraph registers four fixtures, prefixed so they never shadow your own session or graph:
def test_feed(seedgraph_graph):
graph = seedgraph_graph(User, post=2) # fresh in-memory SQLite, FK enforced, tables created on demand
assert len(graph.posts) == 6
async def test_feed_async(seedgraph_agraph):
graph = await seedgraph_agraph(User, post=2)
seedgraph_session and seedgraph_asession expose the sessions behind them. To seed your own database, call seed() on your own session.
Why¶
Every Python team that seeds a relational test database eventually hand-rolls the same plumbing: generate rows, stage commits so primary keys exist, chase those keys into FK columns, repeat for every relationship, and hope the graph stays consistent.
Here is what the alternatives give you, measured on 27 September 2026 by the scripts of seedgraph-comparisons, which anyone can rerun:
| Tool | What you get |
|---|---|
polyfactory 3.3 |
Builds related objects, and SQLAlchemy puts the right keys in the FK columns when it writes them. By default it draws every primary key at random between 0 and 9999, so a test that writes a few dozen rows fails at random with IntegrityError: UNIQUE constraint failed: about a third of runs at 50 posts, three in four at 100, every run at 200. One line fixes it: __set_primary_key__ = False. A shape like 3 users × 2 posts × 5 comments comes out exact when the factory calls are nested from the top (UserFactory.build(posts=[...])). |
faker-sqlalchemy 0.10 (last release August 2022, requires SQLAlchemy < 2.0) |
With generate_related=True, RecursionError on any two-way relationship (backref) and on a self-referential FK; a foreign key passed in overrides is silently replaced by a newly generated parent. |
sqlalchemyseed 2.6 |
Writes data you already have (JSON, YAML, CSV) through your models, nested relationships included. It doesn't generate values. |
sqlseed 0.2 |
Fills an existing SQLite or PostgreSQL database from its schema, not from your models, one row count per table. A graph shape is reachable indirectly: with the coverage strategy, 6 posts over 3 users gives exactly 2 each, so you work out the totals yourself. |
sowdb 0.3 |
PostgreSQL only, from the schema, one row count per table; each foreign key draws a random parent, so 6 posts over 3 users came out as 2, 2, 2 in one run out of ten. |
The failure you meet first, once a test writes enough rows:
import pytest
from polyfactory.factories.sqlalchemy_factory import SQLAlchemyFactory
from sqlalchemy.exc import IntegrityError
class PostFactory(SQLAlchemyFactory[Post]):
__set_relationships__ = True
factory_engine = create_engine("sqlite://")
Base.metadata.create_all(factory_engine)
with Session(factory_engine) as factory_session, pytest.raises(IntegrityError, match="UNIQUE constraint failed: users.id"):
factory_session.add_all(PostFactory.batch(1000)) # 1000 random ids between 0 and 9999 collide
factory_session.flush()
graph = seed(session, User, user=250, post=4) # the database assigns the keys: nothing to collide
assert len(graph.posts) == 1000
polyfactory avoids it with __set_primary_key__ = False on the factory; seedgraph needs no setting.
"But other libraries do this too, don't they?"¶
Some of it, yes: polyfactory builds the same graph once its primary keys are switched off and its calls are nested from the top. What seedgraph adds:
1. The shape in one call. seed(session, User, user=3, post=2, post__comment=5) replaces the nested factory calls. A link that needs a parent takes it from the shape, from parents, or from one generated parent shared by every object that needs it.
2. A verified exit contract. seed() flushes the graph, lets the database assign the keys, then walks every link and raises IncoherentGraphError if a foreign key disagrees with the row it points at.
3. Coexistence with a populated database, with no setting. The database always assigns the keys, so seeding on top of existing rows never collides on ids, never desynchronises a PostgreSQL sequence, and stays safe when two sessions seed the same tables at once: a unique value the other session commits meanwhile is regenerated. Unique columns are checked against the rows already there before anything is written.
4. Determinism wired into pytest. A new session replays the same values from the same seed; consecutive calls on one session continue the sequence instead of repeating it — inside a two-line fixture.
Guarantees and the tests that prove them¶
| Guarantee | Test |
|---|---|
| Every link of the returned graph is verified after flush | test_verification.py::test_seed_returns_a_graph_already_written_with_real_keys, ::test_verify_graph_names_the_link_whose_foreign_key_disagrees |
| Seeding on top of existing rows keeps PostgreSQL sequences intact | test_postgres.py::test_the_application_still_inserts_after_a_seed_on_top_of_its_rows |
| Two sessions seeding the same tables at once do not collide, unique values included | test_postgres.py::test_two_sessions_seeding_the_same_tables_at_once_do_not_collide, ::test_a_unique_value_another_session_commits_mid_seed_is_regenerated |
| Unique columns skip values already in the database | test_unique.py::test_a_new_session_on_a_populated_database_skips_the_values_already_taken, ::test_postgres_rows_from_an_earlier_run_do_not_block_a_new_seed |
| Same seed, same values; consecutive calls do not repeat | test_generators.py::test_a_new_session_replays_the_same_values, ::test_two_seeds_in_one_session_continue_the_same_faker_sequence |
| Generated values fit the column type (enum, length, precision, arrays) | test_types.py |
| Many-to-many shapes build new objects per parent; existing objects are shared | test_many_to_many.py::test_a_many_to_many_count_builds_new_objects_for_each_parent, ::test_existing_objects_passed_as_parents_are_shared_by_every_generated_object |
| Natural and composite primary keys are generated and never collide | test_verification.py::test_a_natural_text_key_is_generated, ::test_natural_keys_skip_the_ones_already_in_the_database, ::test_a_composite_integer_key_and_its_composite_foreign_key_are_generated |
| Multi-column unique constraints hold | test_unique.py::test_many_rows_under_one_parent_keep_a_multi_column_constraint, ::test_a_second_session_keeps_a_multi_column_constraint |
| Existing rows serve as parents; missing required parents are generated once | test_parents.py::test_a_parent_already_in_the_database_is_linked_and_left_out_of_the_graph, ::test_a_child_seeded_alone_gets_one_generated_parent_shared_by_all |
Plugin fixtures live beside a project's own session and graph |
test_fixtures.py::test_the_prefixed_fixtures_live_beside_a_project_own_session_and_graph |
Limits¶
seed()flushes the session. It flushes the objects already pending before building the graph, so the unique checks see them; the keys come from the database, and the graph is no longer pending when it returns.- A shape key towards a parent is refused, as parents are linked or generated on their own; pass existing ones in
parents. View-only relationships are refused too, since nothing would be written. - Required columns of uncovered types (JSON, custom
TypeDecorator, arrays of those) raiseUnsupportedPlaceholderError; declare a generator for them. Nullable ones are left empty. - A multi-column unique constraint whose generated columns are only booleans or enums is left to the database. For the others, one generated column is kept unique on its own, which is stricter than the constraint.
- An association class whose primary key combines its two foreign keys holds one row per parent pair: the generated parent is shared, so two rows under the same parent collide. Seed one per parent, or use a many-to-many relationship.
- A concurrent seed can wait for the other session. When two sessions generate the same unique value, the database holds the second one until the first commits or rolls back; seedgraph then regenerates the value if it was taken. On SQLite the write is not retried, since its drivers open no transaction before a SAVEPOINT: the second session gets the
IntegrityError. - A loop of required links between tables, or a required link to its own table, cannot be generated; pass one side in
parents. - Determinism holds for a given Faker version. Faker may change its data between releases.
Design principles¶
- Model-first, not schema-first. Works from your SQLAlchemy ORM models and relationships.
- Referential consistency is verified, not hoped for. The database assigns the keys, seedgraph checks every link afterwards.
- Shared parents are the point. Realistic data shares parents (one author, many posts). One object per FK is not a graph.
- Deterministic. A new session with the same calls produces the same graph.
- Self-references and mutually referencing tables are normal.
Category.parentand tables pointing at each other are supported; only unsatisfiable loops of required links are refused. - Stop generating at the boundary. Existing rows are usable as parents; only missing parents get generated.
Roadmap¶
- [x] Shape API (
relation=n, nesting, shared parents) - [x] Custom field generators (Faker under the hood)
- [x] Overriding specific attributes on generated objects
- [x] pytest fixture helpers
- [x] Self-referential and cyclic FKs
- [x] Async sessions support
- [x] Database-assigned keys and post-flush verification, PostgreSQL in the test suite
- [x] Type-valid values, uniqueness against existing rows, existing and generated parents
- [x] Many-to-many shapes, generated natural keys, multi-column uniqueness, arrays
- [x] Publication on PyPI
License¶
MIT