Details
-
New Feature
-
Resolution: Unresolved
-
Major
-
18.8.0
-
None
-
Unknown
-
Description
Problem
XWiki has extensive automated coverage of how the UI behaves and almost none of how it looks. Nothing in the build fails when a skin change, a color theme change or a dependency upgrade moves, resizes or restyles an element. Those regressions are found by hand, late, or not at all.
As an example, BlockNote places each block's drag handle at a pixel offset hardcoded per block type and derived from its own default theme (XWIKI-24239). Under XWiki's skin those offsets are wrong, and we expect to keep hitting such cases, both while adapting the editor and on every upgrade of it. We have no way to express "the handle is aligned with its block" as an assertion, and no way to notice that an upgrade silently broke that alignment.
Proposed solution
Compare rendered pixels against a reference image checked into the repository, as a facility of xwiki-platform-test-docker so that any Selenium test can use it. Screenshots are taken with Selenium's WebElement API and compared with image-comparison, which runs inside the build and needs no external service.
Tests get a ScreenshotComparator parameter:
@Test
void sideMenuIsAlignedOnLargeHeadings(TestUtils setup, TestReference testReference,
ScreenshotComparator screenshots) throws Exception
{
setup.createPage(testReference, LARGE_HEADINGS_CONTENT);
BlockNoteRichTextArea textArea = editInplace();
WebElement content = new InplaceEditablePage().getContentContainer();
textArea.hoverBlock(0);
screenshots.assertMatches("heading1", content);
}
References live at src/test/resources/screenshots/<Class>/<method>/<browser>/<name>.png. A missing reference fails the test with a message naming both the screenshot just taken and the path it should be copied to, so the first run of a new test produces the candidate reference and the developer only has to look at it and accept it. On a mismatch the failure reports the percentage of differing pixels and points at a generated diff image.
A working implementation already exists in the BlockNote module and is what the above is based on: https://github.com/mehanix/xwiki-platform/blob/3d9fa71c623713058137d3c687d17d1064fdd375/xwiki-platform-core/xwiki-platform-blocknote/xwiki-platform-blocknote-test/xwiki-platform-blocknote-test-docker/src/test/it/org/xwiki/blocknote/test/ui/ScreenshotComparator.java
What has to be built
Move ScreenshotComparator into xwiki-platform-test-docker and have @UITest register its parameter resolver, so a test only declares the parameter.
Screenshot an element, plus a margin, by cropping the page. The naive approach, screenshotting the element itself, fails for floating UI drawn outside it: BlockNote's side menu appears to the left of the hovered block, so the prototype had to screenshot the nearest ancestor containing both. That ancestor is the full content area, 1220px wide, which makes the test sensitive to every unrelated change inside it, when the test is only about a handle's alignment. Instead, screenshot the page and crop to the element's bounding rect expanded by a caller-supplied inset:
void assertMatches(String name, WebElement element, Insets margin)
The smaller the captured region, the fewer unrelated changes can break the test, which is what makes these tests maintainable at all.
Keep reference images for Firefox only, and skip the assertion (not the whole test) on any other browser. Two reasons. Firefox is the default on CI, so a second reference set would double the maintenance for assertions that mostly run nowhere. And on arm64 the framework substitutes selenium/standalone-chromium for selenium/standalone-chrome; that image doesn't ship fonts-liberation, so sans-serif resolves to Noto Sans, text is wider and paragraphs wrap differently. Running the prototype on Apple Silicon, 6 of 7 Chrome tests fail with 0.1% to 2.8% differing pixels while the alignment under test is in fact correct. Firefox passes 7/7 there: its image is multi-arch, so one reference set covers x86_64 and arm64. Adding Chrome later means first making the tests independent of system fonts.
Make references cheap to update, e.g. -Dxwiki.test.ui.screenshots.update=true writing the generated images straight over the references. Copying files by hand out of target/screenshots is tolerable for one screenshot and painful for twenty, and anything painful here will push people towards weakening the assertions instead of updating them.
Let @UITest configure the viewport size. It is currently around 800px tall and differs between browsers, so tests are written to avoid scrolling. Scrolling in itself should be fine; the case that breaks is hovering an element and then capturing it, because scrolling the target into view can move the hovered element and drop the hover. A taller viewport removes the need to scroll for most tests.
Implemented as screenWidth and screenHeight on @UITest, overridable with -Dxwiki.test.ui.screenWidth and -Dxwiki.test.ui.screenHeight. Note that the reference screenshots depend on the screen width but not on its height, so a test can ask for a taller screen without regenerating them.
Document it on dev.xwiki.org: how to write such a test, how to generate a reference and how to regenerate them all.
Open question: browser version drift
@UITest uses the latest browser image tag and CI pulls it daily. A browser release that nudges text rendering therefore invalidates every reference at once, on every branch, with no change on our side. This is the main recurring cost of the approach and we should decide up front how to absorb it:
- pin browserTag for tests that compare screenshots, and update it deliberately;
- allow a small percentage of differing pixels (setAllowingPercentOfDifferentPixels(), currently 0%), accepting that small genuine regressions pass;
- or treat regeneration as routine and make it a documented one-command operation.
These are not exclusive. Opinions welcome.
Proposal and discussion: https://forum.xwiki.org/t/proposal-visual-regression-testing-in-xwiki-platform/18940