---
title: "Advancing Comprehension of Quantum Application Outputs: A Visualization Technique"
authors:
  - Priyabrata Senapati
  - Tushar M. Athawale
  - David Pugmire
  - Qiang Guan
abstract: Noise in quantum computers presents a challenge for the users of quantum computing despite the rapid progress we have seen in the past few years in building quantum computers. Existing works have addressed the noise in quantum computers using a variety of mitigation techniques since error correction requires a large number of qubits which is infeasible at present. One of the consequences of quantum computing noise is that users are unable to reproduce similar output from the same quantum computer at different times, let alone from various quantum computers. In this work, we have made initial attempts to visualize quantum basis states for all the circuits that were used in quantum machine learning from various quantum computers and noise-free quantum simulators. We have opened up a pathway for further research into this field where we will be able to isolate noisy states from non-noisy states leading to efficient error mitigation. This is where our work provides an important step in the direction of efficient error mitigation. Our work also provides a ground for quantum noise visualization in the case of large numbers of qubits.
summaryType: survey
sourceStatus: null
sources:
  - https://doi.org/10.1145/3588983.3596689
---

# Advancing Comprehension of Quantum Application Outputs: A Visualization Technique

[Read the original paper](https://doi.org/10.1145/3588983.3596689).

## Background and motivation

This short paper investigates how visualization can help researchers inspect the output variation of a quantum machine learning application across many circuit executions.
The application runs on seven qubits, so every measured output is a probability distribution over $2^7 = 128$ computational basis states.
The authors combine functional box plots of these output distributions with heat maps comparing the behavior of individual basis states across the collected executions.
Their intended use is exploratory diagnosis: find states whose distributions warrant further investigation before attempting to separate useful computation from noise.

The underlying problem of unreliable quantum outputs is established rather than new.
In the hardware setting of the 2023 study, noise can make results vary across devices and across repeated use of one device, while the qubit overhead of quantum error correction motivates work on error mitigation.
The related work discusses relaxation time $T_1$, dephasing time $T_2$, CNOT error, and readout error as factors affecting computational reliability.
It also cites QISMET for mitigating temporal noise in variational algorithms and VACSEN for presenting quantum noise through coordinated interactive views.
The gap identified here is narrower: relatively little work visualizes cumulative output behavior across many quantum machine learning jobs, each comprising many circuits.
The paper therefore shifts attention from individual error sources and device characteristics to the distributions produced by the application as a whole.

## Application and data collection

The case study classifies MNIST images of digits 3 and 6 using a model built on TorchQuantum.
The authors train on 2,000 images in a noisy simulator that emulates the target NISQ devices, because access restrictions and IBMQ fair-share scheduling constrain hardware-based training.
Their schedule trains the model once each week and tests it on real machines during the remaining six days.
The reported test devices are IBM Nairobi and IBM Perth, with 100 test images, a batch size of 100, a learning rate of 0.005, and the default circuit-transpilation optimization level.
After amplitude encoding, the circuit applies rotation gates and measurement gates, and qubit-wise probabilities are used to calculate the training loss.
The paper does not provide a detailed circuit diagram or a quantitative comparison of classification accuracy.

A noise-free simulator serves an additional screening role.
When the authors discover that some test images are misclassified even in ideal simulation, they exclude those images and test the remaining images on hardware.
This procedure reduces one source of confusion between model errors and hardware effects, but it does not provide a ground-truth labeling of noisy basis states.
They collect output distributions from all circuits in the hardware jobs, yielding thousands of rows of basis-state probabilities.
Each row corresponds to a circuit output, while a basis-state column records how that state's probability changes across the collected outputs.

The image-encoding description contains an unresolved numerical inconsistency.
The paper says images are compressed from $28 \times 28$ to $12 \times 12$ pixels and then says that 124 pixel values are encoded into the amplitudes of 128 basis states.
Since $12 \times 12 = 144$, the stated dimensions and pixel count do not specify a reproducible mapping without further explanation.
The paper does not explain any additional selection or compression step that reconciles them.

## Visual representations

Figures 1a and 2a summarize the collected output distributions for Nairobi and Perth using functional box plots.
The horizontal axis lists computational basis-state indices and the vertical axis gives probability.
A yellow median curve shows the typical distribution, the dark-blue inner band represents the central 50 percent region, and the cyan outer band represents the non-outlying region.
Figure 2a also overlays outlying curves, identified by a red outlier entry in its legend.
The curves and envelopes reveal which state probabilities remain relatively small and which vary strongly across the aggregated outputs.
The paper names these statistical regions but does not explain the functional ordering or outlier-detection calculation in detail.

Figures 1b and 2b compare every pair of basis-state distributions using what the paper calls KL-distance.
Both matrix axes enumerate the 128 basis states, and color ranges from dark purple for smaller values to green and yellow for larger values.
A bright horizontal or vertical stripe identifies a state whose distribution differs from many other states' distributions.
These matrices compare basis states within a device's collected outputs; they are not direct heat maps of hardware error relative to an ideal output distribution.
The paper does not specify the estimation, normalization, or zero-probability handling used to calculate the plotted divergences.
It presents static figures and does not describe interaction mechanisms or an implemented interactive analysis interface.



The retained image reproduces the two panels of Figure 1 from the paper.
The left panel shows aggregated probabilities and their spread; the right panel makes the prominent stripe at the axis label 49 visible.
These are visualizations of collected application outputs, rather than schematic illustrations of a proposed workflow.

## Findings and strength of the evidence

The demonstration shows different probability envelopes and matrix patterns for the two devices.
The Perth figure was generated from substantially more jobs than the Nairobi figure, so the displayed differences are not a controlled comparison with equal sample sizes.
In the Nairobi matrix, the authors single out a bright stripe as a particularly dissimilar basis-state distribution.
The plot labels this stripe 49, while the accompanying discussion calls it state 49 and then state 50.
The visual evidence establishes the unusual stripe but does not resolve that indexing inconsistency.

The authors hypothesize that states with distinctive distributions and high KL-distance may be dominant contributors to the intended computation, while other excited states may reflect noise.
They explicitly state that this hypothesis requires further study.
The figures therefore identify candidates for investigation, not verified classes of computational and noise-induced states.
A low or high observed probability is not independently established here as evidence that a state is irrelevant or useful to the algorithm.
The reported evaluation consists of the hardware case study and qualitative interpretation of its plots; it does not include a user study, validated noise-state labels, a measured error-mitigation improvement, or a comparison against another visualization method.

## Contributions, limitations, and future work

The contribution is an application-level visualization approach that combines an overview of many output distributions with pairwise comparisons of basis-state behavior.
The functional box plots convey typical values and variation across circuit outputs, while the matrices expose states that differ from many others.
Their use on a real QML workload from two quantum devices establishes a concrete exploratory example, but the paper does not introduce a new error-correction or error-mitigation algorithm.

The scope of the evidence is one seven-qubit application with two MNIST classes and two IBM devices.
Aggregating different input circuits and jobs can expose variation, but the presented analysis does not isolate how much comes from the intended computation, particular noise mechanisms, or changes over time.
The missing divergence-estimation details, inconsistent encoding description, and unequal job counts limit reproducibility and interpretation.
These are reporting and evidential limits of the study, rather than demonstrated failures of the visual encodings.

The authors propose further work on distinguishing basis states involved in computation from states activated by noise, with the longer-term aim of suppressing noise and improving mitigation.
They also motivate visualization for much larger quantum computers, but evaluate only 128 basis states.
Scaling remains an untested issue: enumerating all basis states requires $2^n$ entries for $n$ qubits, and a full pairwise state matrix requires $2^n \times 2^n$ cells.
Consequently, the study supports the use of these views for exploratory analysis at its demonstrated scale, while leaving state attribution, mitigation effectiveness, and large-system visualization for future investigation.
