Methodology · version 1.0
Understand the evidence behind your score.
The Release Visibility Check is a self-assessment of how your team finds and uses test evidence. It describes your answers for your chosen scope. It does not approve a release, predict incidents, or certify your organization.
What contributes to the score?
Only the ten practice questions count. Repository counts, platforms, company size and optional research details do not change your score. A small team and a large organization can both score 100.
| Area | Maximum points |
|---|---|
| Results visibility | 20 |
| Evidence completeness | 25 |
| Evidence traceability | 20 |
| Failure follow-up | 15 |
| Release decisions | 20 |
Each answer has a level from 0 to 4. Its contribution is the question weight multiplied by the level, divided by four. We add the contributions, then round the total to the nearest whole number, with halves rounded up.
For example, level 2 for every question except level 1 for the missing-evidence question gives 46.25 points, displayed as 46/100.
Questions and answer levels
Choose the highest description that is consistently true. You can use your own tools and processes; using QaCockpit never adds points.
When a release decision is needed, how do people get the relevant test results? Weight: 10
Include all projects in your chosen scope. Consider whether people must open each project or ask teams separately to build the summary.
- We cannot reliably find the results needed for the decision. Level 0
- We ask people and search across messages, files, or pipeline pages. Level 1
- We use a known list of sources, then assemble the results for each decision. Level 2
- We have a shared summary of the required sources, updated before the decision. Level 3
- Our shared summary updates as results arrive and clearly shows its scope. Level 4
How much detail is available from the automated tests used in release decisions? Weight: 10
Required sources are the suites or pipelines you expect to review. Can someone move from the summary to the exact failed test and its available diagnostic evidence?
- We have no usable automated test results. Level 0
- We mostly see whether a pipeline passed or failed. Level 1
- We can find individual test outcomes for some of the required sources. Level 2
- We can find individual test outcomes for all required sources. Level 3
- Those outcomes also link directly to the available failure details and diagnostic evidence. Level 4
How do you notice that expected tests or test reports are missing? Weight: 15
Expected scope means the sources and tests you agreed should report. It is different from code coverage or requirements coverage.
- We have no defined check for missing tests or reports. Level 0
- Someone may notice an unusual count or a missing report by chance. Level 1
- We compare the received results with a maintained checklist before the decision. Level 2
- Automated checks flag missing sources or reports against a defined expected scope. Level 3
- Those checks also flag unexpected reductions in test scope, with changes explicitly reviewed. Level 4
Before a release, how are skipped or disabled tests handled? Weight: 10
If no tests were skipped recently, answer for the process that would detect and handle them.
- They are not distinguished from passed tests, or are left out of the summary. Level 0
- They can be found in individual reports, but are not reviewed routinely. Level 1
- They are visible in the release summary and reviewed before the decision. Level 2
- Excluded tests also have recorded reasons and named owners. Level 3
- Exclusions are time-limited, reviewed, and closed with evidence of restored testing. Level 4
Can you tell whether a test result applies to the version being released? Weight: 10
Use an exact commit, artifact digest, package version, or equivalent identifier. A branch name alone can contain several versions.
- We usually cannot link a result to a specific version. Level 0
- We mostly rely on a branch name or whichever result is latest. Level 1
- We can manually trace the exact run, version, and test environment. Level 2
- These identifiers are included in the release summary for every required source. Level 3
- A mismatch with the intended release version or environment is explicitly flagged. Level 4
Could someone later reconstruct which test evidence supported a release decision? Weight: 10
Consider the period your team needs. Keeping evidence indefinitely is not required.
- We do not keep a record of the evidence used. Level 0
- We keep links to changing latest results or rely on people's memory. Level 1
- We record the exact runs, but do not check how long their evidence remains available. Level 2
- We retain the selected evidence for a defined period that meets our needs. Level 3
- We also check that the retained evidence stays accessible and identifiable throughout that period. Level 4
When a required test fails, how is the follow-up owned? Weight: 10
A named team can be an owner. Consider whether the next action remains clear when work changes hands.
- There is no clear person or team responsible for the next action. Level 0
- We ask around in messages to find someone to investigate. Level 1
- A person or team is assigned, but progress is tracked inconsistently. Level 2
- The owner, current status, and next action are visible to the people involved. Level 3
- Unresolved failures are also reviewed against an agreed response time and escalation path. Level 4
What happens when a failed test passes on a rerun? Weight: 5
Answer for your agreed practice even if this has not happened recently. If reruns do not apply, choose Not applicable.
- The passing rerun replaces or hides the original failure. Level 0
- Both attempts remain available, but we usually move on without review. Level 1
- We review both attempts before deciding whether the failure is resolved. Level 2
- Repeated instability also gets a tracked issue and owner. Level 3
- Unstable tests have reviewed, time-limited handling and evidence is required to close the issue. Level 4
How are the test-evidence criteria for a release decision defined? Weight: 10
Consider the written criteria, not just whether a pipeline has a green status.
- There are no agreed criteria. Level 0
- The criteria depend mainly on who is available or how urgent the release is. Level 1
- We use a checklist, but it leaves required evidence or exceptions unclear. Level 2
- Criteria define the required evidence, blockers, and who can approve an exception. Level 3
- We also review and update those criteria using findings from previous releases. Level 4
What happens if required test evidence fails the agreed release criteria? Weight: 10
A checkpoint can be an enforced pipeline rule or a required, recorded human review.
- Release can proceed without a defined check or recorded exception. Level 0
- Someone may stop the release, but the check is informal. Level 1
- A required checkpoint flags the problem, but exceptions are not consistently recorded. Level 2
- Release is held until criteria are met or an authorized exception is recorded. Level 3
- The checkpoint also catches missing or outdated evidence, and exceptions have follow-up owners. Level 4
How to read the result
| Score | Interpretation |
|---|---|
| 0–24 | Limited visibility |
| 25–49 | Partial visibility |
| 50–74 | Established practices |
| 75–100 | Consistent practices |
These are QaCockpit's initial question weights and thresholds, not an industry benchmark or a validated measure of release risk. Small differences between scores should not be treated as meaningful differences in performance.
Unknown, different, not applicable, or skipped
These answers are not given zero points. With fewer than ten scored answers, you see an incomplete result and a possible range based on the unanswered weights. We do not scale your other answers up to 100. With no scored answers, we do not show a score or range.
An area score is available only when both of its questions are answered. Optional explanations are never scored. If a question does not fit, use the relevant option and tell us how to improve it.
Why a high score can still show a gap
Low levels for detecting missing evidence, matching release versions, or applying release checkpoints stay visible even when the total is high. For example, the highest level everywhere except no missing-evidence checks scores 85, with that gap explicitly highlighted.
How we select your next steps
We suggest up to three actions from levels 0–2. The critical gaps above come first, then the largest unused weighted contribution, with question order breaking ties. We show at most one action per area. An incomplete result first suggests gathering the missing information.
If your answers describe consistent practices, we suggest checking that they continue to work. Recommendations can be applied without QaCockpit. Product links depend on the platform you selected and the capabilities of the relevant edition.
Your answers and privacy
The check runs in your browser. There is no account or email gate. Sharing answers for research is a separate, optional action with a preview and a deletion key. Signing up for the Azure launch announcement is independent of your answers.
Read the data handling details or discuss your process with us.
