Testing
Marginalia's tests are integration tests first: most of them start the real Spring XML configuration against a throwaway SQLite database and talk to a fake OpenAI-compatible server. They need no network, no model and no API key, and they run in CI on every push and pull request. This page explains the test infrastructure and how to add tests.
All test code is in marginalia/src/test/java/com/github/enerccio/marginalia/; paths below are relative to it.
Running the tests
From marginalia/:
mvn test # all tests
mvn test -Dtest=LorebookActivationTest # one class
mvn test -Dtest='Backup*Test' # a pattern
mvn test -Dtest='MacroRenderingTest#name' # one method
Surefire is configured in pom.xml with:
Reports are written to target/surefire-reports/; the CI job uploads them as an artifact when a test fails.
In the IDE
Tests run from IntelliJ IDEA like any JUnit test, with two conditions:
-
the classes must be compiled with AspectJ (see Building from source), otherwise UI classes and
@Configurablebeans miss their dependencies; -
add
-XX:+EnableDynamicAgentLoading -Djdk.attach.allowAttachSelf=trueto the VM options of the JUnit run configuration template if you runRuntimeInstrumentationTest.
Without marginalia.test.home the tests fall back to target/test-home relative to the working directory, so run
them with marginalia/ as the working directory (IntelliJ's default for the module).
Libraries
There is no mocking library. Services are tested through the real Spring beans and the real database; the only fake is the model server.
What is covered
Not covered by automated tests: the Vaadin UI (apart from ResourcesPartTest and BranchSearchTest), the bundled plugins, and the desktop launcher
(the CI smoke test only checks that the packaged app starts).
Test bases
classDiagram
MarginaliaTestBase <|-- GenerationTestBase
MarginaliaTestBase <|-- OwnedCrudContract
OwnedCrudContract <|-- ExtendableCrudContract
MarginaliaTestBase <|-- BackupTestBase
GenerationTestBase <|-- LorebookActivationTest
ExtendableCrudContract <|-- ProtocolCrudTest
TemplateTestBase <|-- MacroRenderingTest
class MarginaliaTestBase {
+llm
+createUser()
+login()
+createAI()
}
class GenerationTestBase {
+ai
+protocol
+newManuscript()
+generate()
+onEvent()
}
class TemplateTestBase {
+context
+data
+render()
}TemplateTestBase doesn't start Spring. Pick the lightest base that works:
MarginaliaTestBase
test/MarginaliaTestBase starts the application context from
src/test/resources/META-INF/spring/test-application-config.xml. That file imports the production
application-config.xml and adds TestContextPostProcessor, which
- removes the OSGi framework (
OsgiServiceImpl) and the instrumentation agent (RuntimeInstrumentationInitializer), - points
Configurationto a new foldertarget/test-home/ctx-<random>/instead of~/.marginalia.
Everything else is the production wiring: the services, Hibernate, Flyway migrations, the session-scoped current user. The base class gives you:
Each test method runs in its own mock HTTP request and session (@SpringJUnitWebConfig), so session-scoped beans
start empty: call login() before using a service that needs the current user.
The shared database
Spring caches the application context, and with it the database, for the whole test run. All test classes that use
MarginaliaTestBase share one SQLite database. Write tests so that they don't depend on what else is in it:
- create the data a test needs in the test, with
uniqueName(...)for anything looked up by name; - assert on the entities you created (
contains,doesNotContain), not on totals (hasSize); - operations that work on the whole database (cleanup, database backups) must only assert on their own data - see
CleanupServiceTest.
When a test really needs an empty database or changes global state, mark it with @DirtiesContext. The next test
then gets a new context with a new folder and database:
A new context costs a few seconds (Spring, Hibernate, Flyway), so use it sparingly.
The application folder
Each context gets its own folder under target/test-home/ with the database, data, backup and extension folders.
They are not deleted after the run, so you can open a test database with any SQLite tool after a failure. mvn clean
or deleting target/test-home removes them.
GenerationTestBase
test/GenerationTestBase runs real generations through every pipeline step (see
Generation pipeline). Before each test it logs in a new user, creates an AI pointing to
llm, saves a chat-completion protocol and sets the mock's fallback answer to "generated".
test/GenerationRun is the GenerationListener that stands in for the story editor. It records the streamed
response and reasoning, errors (getErrors()) and simple errors (getSimpleErrors(), the messages the UI shows in a
notification), warnings (getWarnings()), the created part (getMessage()), and the outcome (COMPLETED or CANCELLED). Questions the
pipeline asks the user (askQuestion) are answered yes.
class MyGenerationTest extends GenerationTestBase {
@Test
void instructionsReachTheModel() throws Exception {
Manuscript book = manuscriptService.save(newManuscript());
llm.reply("The lighthouse was dark.");
GenerationRun run = generate(book, "The keeper climbs the stairs.");
assertThat(run.await(TIMEOUT)).isEqualTo(GenerationRun.Outcome.COMPLETED);
assertThat(run.getErrors()).isEmpty();
assertThat(run.getResponse()).isEqualTo("The lighthouse was dark.");
assertThat(llm.lastCompletionRequest().getLastMessage().content())
.contains("The keeper climbs the stairs.");
}
}
To look at the pipeline state in the middle of a generation, capture it in an event listener. LorebookActivationTest
reads the activated entries this way:
AtomicReference<List<LorebookEntry>> activated = new AtomicReference<>();
onEvent(Events.AFTER_PROCESS_LOREBOOK, book,
e -> activated.set(e.getProperty(GenerationProperties.ACTIVATED_LOREBOOK_ENTRIES)));
generate(book, "...");
Listeners are global (StoryGenerationService.addEventListener); onEvent filters by book and always calls
chain.next(), so a failing assertion inside the action doesn't stop the generation. Assert after generate returns,
not inside the listener - an exception in the listener only ends up in the log.
CRUD contracts
Every entity that belongs to a user has a CRUD test built from crud/OwnedCrudContract. The contract contains the
tests; a subclass only says how to make and change an entity:
The inherited tests check identity and owner on create, find by id / uuid / for user, update (creation date kept,
modification date moved), soft delete (hidden from all find*ForUser and findAll* methods), hard delete, isolation
between users, and that saving as another user keeps the owner.
Entities extending ExtendableEntity use crud/ExtendableCrudContract, which adds the persistence of plugin
attributes (getAttributes()): nested JSON, changes and removals, unrelated updates, and saveWithoutEvent
(used by restore). Put at least one @ExtendedAttribute field into newEntity() and modify(), so the inherited
tests also cover the extended fields. crud/ProtocolCrudTest is a short, complete example.
When you add an entity (see Database & migrations), add its <Entity>CrudTest.
TemplateTestBase
templates/TemplateTestBase creates a TemplateServiceImpl directly, without Spring, with a deterministic context:
the clock is fixed at Sunday, 15 March 2026, 14:30:45 UTC, random macros use a seeded Random, and the template data
describes a scene in the middle of a story (POV Alice, characters Alice, Bob, Carol...). One data object is
shared by all render(...) calls of a test, the way lorebook entries of one generation share their variables.
MacroLorebookFixtureTest renders src/test/resources/macro-test.json, a SillyTavern lorebook whose entries
document their expected output. When you change macro behaviour, update the fixture and the expectations together.
The fake LLM server
test/llm/MockLLMServer is a small OpenAI-compatible server on 127.0.0.1 with a random port, built on the JDK's
HttpServer. One shared server lives for the whole test JVM (MockLLMServer.shared()); each test gets its own
scenario with its own base URL, http://127.0.0.1:<port>/<scenario-id>/v1. A provider created with
createAI() points to the test's scenario, so tests never see each other's requests or answers.
It answers:
Programming answers
Chat completions are answered in this order:
- queued responses -
llm.enqueue(...)orllm.reply("text"), one per request; - the responder function -
llm.respondWith(request -> ...), if set and it returns a response; - the fallback -
llm.fallback(...); - otherwise HTTP 500 no response programmed.
MockLLMResponse builds one answer:
MockLLMResponse.text("Hello world") // one content chunk
MockLLMResponse.chunks("Hel", "lo").chunkDelay(ofMillis(50))
MockLLMResponse.split(longText, 20) // chunks of 20 characters
MockLLMResponse.text("answer").withReasoning("thinking") // reasoning chunks first
MockLLMResponse.text("answer").withReasoning("...").reasoningField("reasoning") // instead of reasoning_content
MockLLMResponse.text("cut").finishReason("length")
MockLLMResponse.error(429, "slow down") // OpenAI-style error body
MockLLMResponse.chunks("a", "b", "c").disconnectAfter(2) // stream ends abruptly
MockLLMResponse.text("late").waitFor(latch) // held until the latch is released
waitFor(latch) is the tool for "while generating" states: the request is received, the pipeline waits for the
stream, and the test can stop the generation or inspect the state before releasing the latch.
Retries
The openai-java client retries 408, 409, 429 and 5xx responses. A queued error response is consumed once per
attempt, so enqueue it several times (or use fallback) when a test expects the error to reach Marginalia.
Inspecting requests
Every request, including model listing, is recorded:
For code that calls an InferenceService directly (without the pipeline), test/InferenceCollector is a callback
that pulls the whole stream and waits for the end.
Expected errors in the log
Tests of failure paths make the code log errors on purpose. To keep them out of the test output (where they look like
real failures) and to assert on them, capture the logger with test/ExpectedLog:
try (ExpectedLog log = ExpectedLog.capture(ExtensionServiceImpl.class)) {
...
assertThat(log.errors()).anyMatch(m -> m.contains("broken"));
}
While the capture is open, the logger's events are recorded and not passed to the console; errors(), warnings()
and atLeast(level) return the messages, entries() the full events with their exceptions.
Instrumentation tests
The instruct tests check the bytecode that makes @Extendable methods extensible (see
Plugin development):
-
ExtendableMethodVisitorTesttransforms fixture classes frominstruct/fixture/withinstruct/Instrumenter- the same transformation the agent applies, without an agent - and checks that the result verifies, keeps the class schema, behaves like the original and exposes arguments and locals to decorators. -
RuntimeInstrumentationTestinstalls the real agent into the test JVM and checks classes loaded both before and after the installation (fixture/agent/). -
ExtensionVerifierTestruns the extension verification on the extensions offixture/verify/VerifyFixtures(decorators that are valid, ask for things that don't exist, or can't be followed) against the decoratedVerifyTarget. Their class files are packed into a JAR on the fly and served by a stand-inBundle. Decorators must be anonymous classes with the context calls inside them, as in a real extension - a decorator that delegates to a helper class would not be checked.
The fixtures are compiled with the same -parameters and preserveAllLocals options as the application; if you
change the compiler settings in pom.xml, these tests tell you whether extensions still see names of arguments and
locals.
Checklist for a new test
-
Name it
<Subject>Testand put it in the package of the area it tests; Surefire runs every*Testclass. -
Extend the lightest base that works (
Test bases ). -
Create your own data with unique names; don't depend on what other tests left in the shared database.
-
Don't call real models or other network services - program
llminstead. -
Assert with AssertJ, with
.as("...")where the failure message would otherwise be unclear. -
Capture expected error logs with
ExpectedLog. -
A new entity gets a CRUD contract test; a new migration is covered by
FlywayMigrationTestautomatically (it runs all migrations and validates the schema against the entities). -
A bug fix comes with a test that fails without the fix, where the area is testable without the UI.
-
Run
mvn testbefore opening a pull request - CI runs the same command.