Skip to content

seedgraph

Seed your SQLAlchemy models as a referentially-consistent graph — one call, shared parents, verified links.

PyPI CI Python License

seedgraph fills your test database with a coherent graph of objects, in a single call. You declare what you want — "3 users, each with 2 posts, each post with 3 comments" — and the library builds the objects, links them, writes them to your session and then verifies every foreign key against the row it points at. Values are realistic and reproducible ("Jose Bishop", not "user-0"), generated by Faker with a fixed seed, valid for each column's type, and unique where the schema says so — even against rows already in the database. You can pin any column, replace how any column is generated, and attach the new graph to rows you already have. It plugs into pytest with no configuration.

Status: alpha — the API may still change before 1.0. Every guarantee below is backed by a named test, on SQLite and on PostgreSQL.

Documentation: jrachid.github.io/seedgraph — recipes for pytest, FastAPI and custom column types, the API reference, and a Claude Code plugin that teaches coding agents to use seedgraph.

Quick start

pip install seedgraph
pip install "seedgraph[async]"  # for seed_async and the async pytest fixtures

Requires Python 3.11+, SQLAlchemy 2.x and Faker 30+. The async extra adds greenlet (through sqlalchemy[asyncio]) and aiosqlite.

Declare a shape from a root model; each key walks a one-to-many or many-to-many relationship, by relationship name or by target class name:

from sqlalchemy import ForeignKey, create_engine
from sqlalchemy.orm import DeclarativeBase, Mapped, Session, mapped_column, relationship

from seedgraph import seed


class Base(DeclarativeBase):
    pass


class User(Base):
    __tablename__ = "users"
    id: Mapped[int] = mapped_column(primary_key=True)
    name: Mapped[str]
    email: Mapped[str] = mapped_column(unique=True)
    posts: Mapped[list["Post"]] = relationship(back_populates="author")


class Post(Base):
    __tablename__ = "posts"
    id: Mapped[int] = mapped_column(primary_key=True)
    title: Mapped[str]
    subtitle: Mapped[str | None]
    author_id: Mapped[int] = mapped_column(ForeignKey("users.id"))
    author: Mapped[User] = relationship(back_populates="posts")
    comments: Mapped[list["Comment"]] = relationship(back_populates="post")


class Comment(Base):
    __tablename__ = "comments"
    id: Mapped[int] = mapped_column(primary_key=True)
    body: Mapped[str]
    post_id: Mapped[int] = mapped_column(ForeignKey("posts.id"))
    post: Mapped[Post] = relationship(back_populates="comments")


engine = create_engine("sqlite://")
Base.metadata.create_all(engine)
session = Session(engine)

graph = seed(session, User, post=2, post__comment=3)   # 3 users by default

assert len(graph.users) == 3
assert len(graph.posts) == 6
assert len(graph.comments) == 18
post = graph.users[0].posts[0]
assert post.author_id == graph.users[0].id          # real key, assigned by the database

seed() returns once the graph is flushed and verified; commit or roll back as your test needs. seed_async(async_session, ...) is its twin for an AsyncSession. The graph exposes every table of the model's metadata as an attribute, empty when the shape built none.

Existing and missing parents

alice = session.get(User, 1)
graph = seed(session, Post, parents=[alice])   # every post's author is alice; alice is not in graph.users

graph = seed(session, Comment)   # one Post and one User are generated, shared by all comments

A link first takes the nearest ancestor of its type in the shape, then the object of that type passed in parents (optional links included). A required link still empty gets one generated parent per type, shared by every object that needs it. Several objects of one type are accepted in parents; a link towards a single parent refuses to choose between them with AmbiguousParentError.

Many-to-many

from sqlalchemy import Column, Table

article_tag = Table(
    "article_tag",
    Base.metadata,
    Column("article_id", ForeignKey("articles.id"), primary_key=True),
    Column("tag_id", ForeignKey("tags.id"), primary_key=True),
)


class Tag(Base):
    __tablename__ = "tags"
    id: Mapped[int] = mapped_column(primary_key=True)
    name: Mapped[str] = mapped_column(unique=True)


class Article(Base):
    __tablename__ = "articles"
    id: Mapped[int] = mapped_column(primary_key=True)
    title: Mapped[str]
    tags: Mapped[list[Tag]] = relationship(secondary=article_tag)


Base.metadata.create_all(engine)
python, sql = Tag(name="python"), Tag(name="sql")
session.add_all([python, sql])

graph = seed(session, Article, article=5, tags=3)                  # 15 new tags, 3 per article
assert len(graph.tags) == 15

graph = seed(session, Article, article=5, parents=[python, sql])   # every article tagged with both existing tags
assert all(article.tags == [python, sql] for article in graph.articles)

A count keeps its one-to-many meaning: new objects for each parent. Objects passed in parents join every many-to-many collection of their type, next to the ones the shape builds. SQLAlchemy writes the association rows itself.

Pinning and generating values

graph = seed(
    session,
    User,
    post=2,
    generators={User: {"name": lambda ctx: ctx.fake.first_name()}},   # replace how a column is generated
    overrides={Post: {"title": "Imposed", "subtitle": None}},        # pin a value, None included
)

ctx.fake is the session's seeded Faker; ctx.column is the column name. An override can also be a callable taking the same context.

pytest

Installing seedgraph registers four fixtures, prefixed so they never shadow your own session or graph:

def test_feed(seedgraph_graph):
    graph = seedgraph_graph(User, post=2)       # fresh in-memory SQLite, FK enforced, tables created on demand
    assert len(graph.posts) == 6

async def test_feed_async(seedgraph_agraph):
    graph = await seedgraph_agraph(User, post=2)

seedgraph_session and seedgraph_asession expose the sessions behind them. To seed your own database, call seed() on your own session.

Why

Every Python team that seeds a relational test database eventually hand-rolls the same plumbing: generate rows, stage commits so primary keys exist, chase those keys into FK columns, repeat for every relationship, and hope the graph stays consistent.

Here is what the alternatives give you, measured on 27 September 2026 by the scripts of seedgraph-comparisons, which anyone can rerun:

Tool What you get
polyfactory 3.3 Builds related objects, and SQLAlchemy puts the right keys in the FK columns when it writes them. By default it draws every primary key at random between 0 and 9999, so a test that writes a few dozen rows fails at random with IntegrityError: UNIQUE constraint failed: about a third of runs at 50 posts, three in four at 100, every run at 200. One line fixes it: __set_primary_key__ = False. A shape like 3 users × 2 posts × 5 comments comes out exact when the factory calls are nested from the top (UserFactory.build(posts=[...])).
faker-sqlalchemy 0.10 (last release August 2022, requires SQLAlchemy < 2.0) With generate_related=True, RecursionError on any two-way relationship (backref) and on a self-referential FK; a foreign key passed in overrides is silently replaced by a newly generated parent.
sqlalchemyseed 2.6 Writes data you already have (JSON, YAML, CSV) through your models, nested relationships included. It doesn't generate values.
sqlseed 0.2 Fills an existing SQLite or PostgreSQL database from its schema, not from your models, one row count per table. A graph shape is reachable indirectly: with the coverage strategy, 6 posts over 3 users gives exactly 2 each, so you work out the totals yourself.
sowdb 0.3 PostgreSQL only, from the schema, one row count per table; each foreign key draws a random parent, so 6 posts over 3 users came out as 2, 2, 2 in one run out of ten.

The failure you meet first, once a test writes enough rows:

import pytest
from polyfactory.factories.sqlalchemy_factory import SQLAlchemyFactory
from sqlalchemy.exc import IntegrityError


class PostFactory(SQLAlchemyFactory[Post]):
    __set_relationships__ = True


factory_engine = create_engine("sqlite://")
Base.metadata.create_all(factory_engine)
with Session(factory_engine) as factory_session, pytest.raises(IntegrityError, match="UNIQUE constraint failed: users.id"):
    factory_session.add_all(PostFactory.batch(1000))   # 1000 random ids between 0 and 9999 collide
    factory_session.flush()

graph = seed(session, User, user=250, post=4)   # the database assigns the keys: nothing to collide
assert len(graph.posts) == 1000

polyfactory avoids it with __set_primary_key__ = False on the factory; seedgraph needs no setting.

"But other libraries do this too, don't they?"

Some of it, yes: polyfactory builds the same graph once its primary keys are switched off and its calls are nested from the top. What seedgraph adds:

1. The shape in one call. seed(session, User, user=3, post=2, post__comment=5) replaces the nested factory calls. A link that needs a parent takes it from the shape, from parents, or from one generated parent shared by every object that needs it.

2. A verified exit contract. seed() flushes the graph, lets the database assign the keys, then walks every link and raises IncoherentGraphError if a foreign key disagrees with the row it points at.

3. Coexistence with a populated database, with no setting. The database always assigns the keys, so seeding on top of existing rows never collides on ids, never desynchronises a PostgreSQL sequence, and stays safe when two sessions seed the same tables at once: a unique value the other session commits meanwhile is regenerated. Unique columns are checked against the rows already there before anything is written.

4. Determinism wired into pytest. A new session replays the same values from the same seed; consecutive calls on one session continue the sequence instead of repeating it — inside a two-line fixture.

Guarantees and the tests that prove them

Guarantee Test
Every link of the returned graph is verified after flush test_verification.py::test_seed_returns_a_graph_already_written_with_real_keys, ::test_verify_graph_names_the_link_whose_foreign_key_disagrees
Seeding on top of existing rows keeps PostgreSQL sequences intact test_postgres.py::test_the_application_still_inserts_after_a_seed_on_top_of_its_rows
Two sessions seeding the same tables at once do not collide, unique values included test_postgres.py::test_two_sessions_seeding_the_same_tables_at_once_do_not_collide, ::test_a_unique_value_another_session_commits_mid_seed_is_regenerated
Unique columns skip values already in the database test_unique.py::test_a_new_session_on_a_populated_database_skips_the_values_already_taken, ::test_postgres_rows_from_an_earlier_run_do_not_block_a_new_seed
Same seed, same values; consecutive calls do not repeat test_generators.py::test_a_new_session_replays_the_same_values, ::test_two_seeds_in_one_session_continue_the_same_faker_sequence
Generated values fit the column type (enum, length, precision, arrays) test_types.py
Many-to-many shapes build new objects per parent; existing objects are shared test_many_to_many.py::test_a_many_to_many_count_builds_new_objects_for_each_parent, ::test_existing_objects_passed_as_parents_are_shared_by_every_generated_object
Natural and composite primary keys are generated and never collide test_verification.py::test_a_natural_text_key_is_generated, ::test_natural_keys_skip_the_ones_already_in_the_database, ::test_a_composite_integer_key_and_its_composite_foreign_key_are_generated
Multi-column unique constraints hold test_unique.py::test_many_rows_under_one_parent_keep_a_multi_column_constraint, ::test_a_second_session_keeps_a_multi_column_constraint
Existing rows serve as parents; missing required parents are generated once test_parents.py::test_a_parent_already_in_the_database_is_linked_and_left_out_of_the_graph, ::test_a_child_seeded_alone_gets_one_generated_parent_shared_by_all
Plugin fixtures live beside a project's own session and graph test_fixtures.py::test_the_prefixed_fixtures_live_beside_a_project_own_session_and_graph

Limits

  • seed() flushes the session. It flushes the objects already pending before building the graph, so the unique checks see them; the keys come from the database, and the graph is no longer pending when it returns.
  • A shape key towards a parent is refused, as parents are linked or generated on their own; pass existing ones in parents. View-only relationships are refused too, since nothing would be written.
  • Required columns of uncovered types (JSON, custom TypeDecorator, arrays of those) raise UnsupportedPlaceholderError; declare a generator for them. Nullable ones are left empty.
  • A multi-column unique constraint whose generated columns are only booleans or enums is left to the database. For the others, one generated column is kept unique on its own, which is stricter than the constraint.
  • An association class whose primary key combines its two foreign keys holds one row per parent pair: the generated parent is shared, so two rows under the same parent collide. Seed one per parent, or use a many-to-many relationship.
  • A concurrent seed can wait for the other session. When two sessions generate the same unique value, the database holds the second one until the first commits or rolls back; seedgraph then regenerates the value if it was taken. On SQLite the write is not retried, since its drivers open no transaction before a SAVEPOINT: the second session gets the IntegrityError.
  • A loop of required links between tables, or a required link to its own table, cannot be generated; pass one side in parents.
  • Determinism holds for a given Faker version. Faker may change its data between releases.

Design principles

  1. Model-first, not schema-first. Works from your SQLAlchemy ORM models and relationships.
  2. Referential consistency is verified, not hoped for. The database assigns the keys, seedgraph checks every link afterwards.
  3. Shared parents are the point. Realistic data shares parents (one author, many posts). One object per FK is not a graph.
  4. Deterministic. A new session with the same calls produces the same graph.
  5. Self-references and mutually referencing tables are normal. Category.parent and tables pointing at each other are supported; only unsatisfiable loops of required links are refused.
  6. Stop generating at the boundary. Existing rows are usable as parents; only missing parents get generated.

Roadmap

  • [x] Shape API (relation=n, nesting, shared parents)
  • [x] Custom field generators (Faker under the hood)
  • [x] Overriding specific attributes on generated objects
  • [x] pytest fixture helpers
  • [x] Self-referential and cyclic FKs
  • [x] Async sessions support
  • [x] Database-assigned keys and post-flush verification, PostgreSQL in the test suite
  • [x] Type-valid values, uniqueness against existing rows, existing and generated parents
  • [x] Many-to-many shapes, generated natural keys, multi-column uniqueness, arrays
  • [x] Publication on PyPI

License

MIT