12-DEC-2024
0004 - srailimaf
Starting in on #2 from above, the first step will be doing some work working on the Unplay side of things and maybe
doing some genericizing.
1000 - srailimaf
I’ve been poking around at some other tools I might want to integrate, and came across insta, which I think would
greatly improve the tests for the UI. More generally, I want to start considering how to refactor the test code. I like
being able to put tests right alongside the code it’s testing, but I think ultimately I’d prefer if the test code were a
little less heavily duplicated.
I found rstest which offers a little extra, and I think I have a plan for how to refactor longterm.
- Keep developing using
#[test]and the builtin runner along side the source. These are ‘in situ’ unit tests, they are intended to simply do whatever is necessary to make whatever assertion is intended. tests/unit/contains an equivalent folder structure, but usesrstestto run the tests, and is more heavily factored and intended to be the place wherein situtests are moved to once the functionality is stablized (and really, once I get tired of the long file size).tests/integration/can contains integration teststests/uifor ui tests viainsta
and so on.
The goal would be to slowly extract to some generic ‘spec’ – especially for integration tests, relying on non-hazel-specific tools to implement the test as much as possible will allow for easy cross-comparison with knonw-good engines. The spec will ideally be engine-agnostic to some extent, so that it should be relatively low maintenance as I keep tweaking stuff in hazel.
I have some loose refactoring that happens in the unit tests as of right now, and coverage is pretty good, but integration is a little weak. I’ve been thinking about other metrics that could be useful with respect to coverage, in particular I’ve been thinking about two in particular:
- Lines / Test Coverage - How many lines does a particular test touch in our code?
- assertions / line - How many assertions are made per line of code?
Ideally we want a test to cover as few lines as possible (tests should be precise), while making as many (good)
assertions as possible per line. The former is probably easy to calculate based on the .lcov information, but the latter
is a little trickier. I suspect that these are the underlying metrics that mutant-style testing reveals. Having a high
assertion-per-line ratio would mean that random changes to code are more likely to be caught by some test, reducing
the class of ‘tests-missing-obvious-logic-error’ mutant, but not the ‘subtle-logical-change’ mutant. The latter is far
more rare than the former. Fewer lines touched per test means that the test is making more targeted assertions, which
should address the latter.
I may take a sidequest at some point to investigate how to calculate these metrics, if nothing else than because I like a good metric.