---
title: Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing
authors:
  - Priyabrata Senapati
  - Qiang Guan
  - David Pugmire
  - Cheng Chang Lu
  - Tushar M. Athawale
abstract: Quantum computing technology holds substantial promise as a reliable computational paradigm. However, current noisy intermediate scale quantum (NISQ) systems, are significantly impacted by noise originating from hardware inconsistencies. This noise causes errors and lowers output fidelity. So we must find which basis states cause errors. However, there are two main challenges in analyzing noise corresponding to basis states. First, the noise distribution data is high dimensional in nature, thereby making its analysis challenging. Second, although functional box plots have been used in the state of the art research to understand such a high dimensional data, they suffer from clutter and occlusion issues because of overplotting. In this study, we introduce an innovative visualization pipeline to address the aforementioned challenges to provide a clear depiction of noisy and less-noisy basis states. Specifically, our proposed visualization pipeline comprises three stages namely, low dimensional embedding, clustering, and violin plot visualization, to reduce visual clutter and effectively analyze high-dimensional noise distribution data. Our analysis uses quantum machine learning (QML) circuits as case study for drawing a distinction between noisy and less noisy basis states.
summaryType: survey
sourceStatus: null
sources:
  - https://doi.org/10.1109/QSW67625.2025.00011
---

# Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing

[Read the original paper](https://doi.org/10.1109/QSW67625.2025.00011)

## Background and motivation

Senapati and colleagues present a visualization pipeline for inspecting variation in computational-basis-state probabilities from quantum machine learning experiments.
A register of $n$ qubits has $2^n$ computational basis states, and repeated circuit measurements estimate the probability of each outcome.
Hardware imperfections, crosstalk, readout errors, and temporal drift can alter these probabilities, while finite sampling also introduces shot noise.
The paper addresses the established problem of understanding this variability in noisy intermediate-scale quantum systems, focusing on which basis states exhibit relatively stable or variable behavior across many circuit executions.
Its contribution is a way to organize and display the resulting distributions, rather than a new quantum algorithm or error-mitigation procedure.

The related work includes QubiCSV for calibration-data management and visualization, VACSEN for hardware and circuit noise awareness, and state representations such as VENUS and BEADS.
The most direct predecessor is the authors' use of functional box plots to summarize quantum application outputs.
These plots overlay central regions, medians, and outliers across basis states, but their dense bands and curves make the behavior of individual states difficult to distinguish.
Figure 2a(i) illustrates this problem for 128 states in a seven-qubit system.
The proposed alternative first groups states with similar distributions and then shows a smaller set of representative distributions, addressing clutter through aggregation rather than attempting to display every high-dimensional observation at once.

## Experimental data and the meaning of variability

The case study uses a seven-qubit classifier for MNIST digits 3, 6, and 9, with hardware runs on IBM Lagos, IBM Nairobi, and IBM Perth.
The authors train with 2,000 images on a noisy simulator and describe a weekly schedule of training on one day followed by hardware testing on the other six days, subject to cloud access and queue delays.
Testing uses batches of 300 images.
Before hardware execution, they exclude images misclassified by a noise-free simulator, so the retained test set has 100% accuracy under that simulator by construction.
This selection is intended to focus the subsequent analysis on execution-related degradation; it is not a reported generalization accuracy for the original, unfiltered test set.

The analysis uses a matrix $X \in \mathbb{R}^{8000 \times 128}$ derived from shot counts.
Each row records the empirical basis-state probability distribution for an image-circuit execution, and unobserved basis states receive zero probability.
Each column therefore contains 8,000 probability observations for one basis state, rather than 8,000 amplitudes of a single quantum state.
The visual analysis compares the distributions of these column values.

An important qualification is that the rows also involve different images within a class.
The introduction explicitly recognizes two sources of variation: noise in quantum execution and the natural differences in image pixel values that change basis-state activations.
Although the methodology describes simulator-based filtering as eliminating image-originating noise, correct classification alone does not establish that different images have identical ideal output distributions.
The plotted variability should therefore be understood as observed variation across the selected executions, which the authors interpret as relative noise, rather than a demonstrated separation of hardware noise from all application-data variation.

## Distribution comparison, projection, and clustering

Figure 1 is a schematic of the processing sequence, from QML data collection through pairwise distribution comparison, dimensionality reduction, clustering, and final visualization.
For each basis-state column, the method constructs a histogram of the observed probabilities and computes pairwise Kullback-Leibler divergence between the histograms.
The resulting $128 \times 128$ matrix is displayed as a heat map, with a dark blue-purple region indicating relatively similar distributions and yellow indicating larger differences.
The heat map compares distribution shapes; a yellow row or column identifies a state unlike many others, not by itself a state with the highest variance.

Multidimensional scaling uses these KL-based dissimilarities to place the 128 basis states in two dimensions.
Each point represents one basis state, and nearby points are intended to have more similar probability distributions.
The two axes are embedding coordinates, not physical qubit positions or individual probabilities.
The method then applies $k$-means to these projected points.
An elbow plot of distortion, defined as the sum of squared distances to the nearest cluster centroid, guides the choice of the number of clusters.
Cluster membership is encoded with categorical point colors and numerical basis-state labels.
This combines established statistical methods into an application-specific analysis workflow rather than introducing a new embedding or clustering algorithm.

## Visual encoding and analysis workflow

The final stage presents violin plots for the distributions associated with the cluster representatives, which the paper calls centroid basis states.
A violin's width depicts estimated density, its vertical extent shows the range and spread of observations, and internal marks summarize the distribution.
These views allow readers to compare concentration, tails, and variability without overlaying all 128 basis-state curves.
The paper labels the vertical axis of its centroid plots as “variance,” but describes the violins as distributions across the raw-data rows; the relevant visual comparison is the relative spread of those distributions, rather than one scalar variance value per violin.

After inspecting the centroid violins, the authors identify clusters with relatively high spread and recolor their member states red in the MDS view.
States in the other clusters appear blue.
This sequence separates two uses of color: categorical colors first show similarity groups, while red and blue subsequently show the authors' relative noise classification.
Additional violins for selected members provide examples of within-cluster distribution similarity.
The implementation is described as a suite of Python scripts accepting raw measurement readouts or aggregated Qiskit counts, with the user pointing the scripts to a results directory.
The paper presents generated plots and does not describe an interactive application with linked brushing, selection controls, or a user-driven noise threshold.



Figure 2 shows actual analysis outputs for class 3 images on IBM Lagos.
The top row contrasts the cluttered functional box plot with the distribution-comparison heat map and MDS layout.
The middle row introduces cluster selection and membership, while the bottom row uses representative violins to assign relative noise categories back to the embedded basis states.

## Findings and contribution

For Lagos class 3, the heat map and MDS view highlight states 30 and 61 as distributionally distinct from much of the remaining data.
Figure 2's elbow plot selects $K=10$.
Violin plots for states 4, 20, 40, and 56 illustrate similar distribution shapes within one cluster, and another example compares states 13, 22, 76, and 95.
The authors identify clusters represented by states 50, 61, and 105 as having greater spread than examples represented by 22 and 117, and use that distinction to color the final scatterplot.
These examples show how distribution comparison and representative inspection support the classification, but they do not constitute a quantitative validation against independently known noise labels.

Figures 3–5 apply the process to class 3 on Nairobi and Perth.
State 8 stands apart in the Nairobi comparison and state 36 in the Perth comparison; both devices use eight clusters in these examples.
For Nairobi, the text identifies centroid states 11, 22, and 80 as relatively variable.
The appendix extends the displayed analysis to classes 6 and 9 on all three devices in Figures 6–11, using eight or nine clusters depending on the case.
The Lagos appendix examples also include enlarged views with the isolated state 30 omitted to make the remaining layout easier to read.
These panels demonstrate that the same workflow can be applied to several device-and-class combinations, rather than establishing that particular basis-state identifiers are intrinsically noisy across workloads or devices.

The main contribution is a compact basis-state-level diagnostic representation combining distribution-sensitive grouping with explicit views of distribution shape.
Compared with the functional box plot example, the representative violins expose variation with fewer overlaid marks, and the colored MDS view carries the resulting group interpretation back to individual states.
The evaluation consists of these hardware-derived case studies and visual comparisons.
The paper does not report a controlled user study, task-completion measurements, or a quantitative clutter metric, so its claims about clearer and faster interpretation rest on the presented examples rather than measured analyst performance.

## Limitations and future directions

The method supplies relative descriptive labels, not causal attribution of an observed effect to a particular physical qubit or error mechanism.
The paper does not specify a general numerical threshold separating noisy from less noisy clusters, and the summary plots depend on histogram construction, embedding, clustering, and representative selection.
It provides no projection-quality metric or sensitivity analysis establishing that the two-dimensional arrangement preserves all important distributional relationships.
Its examples support diagnostic inspection of seven-qubit data, while scaling to larger systems, whose basis-state count grows exponentially, remains future work.

Data collection took nearly 20 months under restrictive backend access, and the authors acknowledge that the processors had become legacy systems as IBM hardware advanced.
The reported state patterns therefore describe the collected historical workloads, not present-day backend capabilities.
Although the discussion proposes using these insights to guide circuit placement or error mitigation, the paper does not demonstrate a placement change, a mitigation algorithm, or a resulting improvement in prediction accuracy.
The conclusion explicitly identifies more direct localization of noisy qubits as an extension.
Further plans include applications beyond QML, support beyond the Qiskit ecosystem, variational applications such as VQE and VQLS, non-variational algorithms such as phase estimation, and larger devices with more complex noise propagation and multi-qubit effects.
