Debug TensorFlow models in stages: first make a small example work in eager execution, then reproduce graph-only behavior, locate the first operation that creates a NaN or infinity, and profile slow training steps before changing hardware or scaling out. This sequence separates correctness problems from execution-mode surprises and performance bottlenecks.
Start with a small, inspectable eager-mode baseline
TensorFlow 2 eager execution makes it easier to step through operations and inspect intermediate values. Reduce the problem to a small, reproducible input, then run the affected model or training step eagerly. Check the input shapes and dtypes, labels, model outputs, loss, and gradients. TensorFlow’s Effective TensorFlow 2 guide and tf.function guide recommend getting code to execute without errors in eager mode before applying tf.function where graph execution is needed.
Once the eager version behaves as expected, restore the execution path that reproduces the problem. That distinction matters: a bug that disappears in eager mode may depend on tracing or graph execution, while a failure in both modes is more likely to be in the data, model logic, or numerical operations.
When behavior changes inside tf.function
Python statements inside a tf.function do not all behave like statements in an ordinary step-by-step Python run. A regular Python print runs when TensorFlow traces the function, so it is useful for seeing when tracing occurs. It does not necessarily report tensor values each time the graph executes. Use tf.print for runtime tensor values.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
@tf.function
def step(x):
print("Tracing step") # Python: runs during tracing
tf.print("x:", x) # TensorFlow op: runs with the graph
return model(x)
If you need to inspect a function step by step, temporarily enable eager execution for functions:
tf.config.run_functions_eagerly(True)
# Run the function and inspect its behavior.
# Restore normal function execution afterward:
tf.config.run_functions_eagerly(False)
Use this as a diagnostic setting, not as a performance fix: it changes how functions run and is intended to make debugging easier. TensorFlow summarizes the trade-off in its tf.function guide: “In general, debugging code is easier in eager mode than inside tf.function.”
Rank #2
- Machine Learning Using TensorFlow Cookbook: Create powerful machine learning algorithms with TensorFlow
- ABIS BOOK
- Packt Publishing
Find the first operation that creates a NaN or infinity
If a loss, activation, gradient, or weight becomes non-finite, inspecting only the final loss tells you that something went wrong, not where. Enable numerical checks to stop when an operation produces NaN or infinity:
tf.debugging.enable_check_numerics()
# Run the operation or training step that reproduces the failure.
This is a focused way to identify the operation that first emits an invalid value. For a few known tensors at a known location, strategically placed tf.print statements may be enough to inspect the inputs and outputs around that operation.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #3
Use Debugger V2 when the source is unclear
When the invalid value’s origin is not obvious, TensorBoard Debugger V2 can provide broader context: execution history, tensor summaries or values, tensor-health information, graph structure, source locations, and stack traces. The Debugger V2 guide advises inserting tf.debugging.experimental.enable_dump_debug_info() early enough to record the relevant program activity. Use the guide’s instructions for setting up and inspecting a debug dump.
Debugger instrumentation adds overhead, and its impact depends on the debug mode, hardware, and workload. Use it to investigate a problem, rather than treating an instrumented run as a performance benchmark.
Rank #4
Fix the invalid operation, not just the symptom
The Debugger V2 tutorial demonstrates a negative infinity caused by taking the logarithm of zero-valued probabilities. In that example, clipping values before the logarithm or using tf.keras.losses.CategoricalCrossentropy are possible remedies. Those are example-specific options, not universal cures: first identify the operation and the invalid input, then choose a correction that matches the model’s intended mathematics.
Profile a slow step before tuning the GPU
A GPU that appears underutilized may be waiting for input, host-side work, or another part of the execution pipeline. Use TensorFlow Profiler through TensorBoard to establish where time is going instead of guessing from a utilization reading. Its overview and trace tools help distinguish device computation from idle time, host-to-device activity, and input-pipeline delays. TensorFlow describes profiling as a way to understand the time and memory consumed by TensorFlow operations and resolve performance bottlenecks in its Profiler guide.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
Check input delivery, then inspect the trace
Start with the input-pipeline analyzer to determine whether the run is input-bound. If input work is blocking the device, inspect the pipeline stages and benchmark the input pipeline independently so that loader changes are not confused with model or backpropagation time. TensorFlow’s tf.data performance guide recommends placing prefetch at the end of the input pipeline to overlap input work with model computation.
If the analyzer does not point to input delivery, use the trace and overview to investigate host and device timing patterns. Diagnose the bottleneck on a single GPU before investigating multi-GPU behavior; scaling up before locating the single-device constraint can obscure the cause rather than solve it. The GPU performance analysis guide covers that diagnostic approach.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Compare training behavior when migrating from TensorFlow 1.x to 2.x
When a migrated pipeline runs but trains differently, compare the quantities that can reveal where behavior first diverges—not just final accuracy. TensorFlow’s migration debugging guide identifies these comparisons:
- Learning rate and model weights.
- Gradient scale.
- Training and validation metrics.
- Intermediate outputs.
Track them over the run and look for the first meaningful difference. This narrows the investigation to the stage where the migrated training behavior begins to depart from the reference.
Recommended Free Tools
Choose the diagnostic by the question you need answered
| Diagnostic choice | Best suited to |
|---|---|
| Eager execution | Step-by-step inspection and establishing a readable baseline. |
Graph execution with tf.function |
Reproducing issues that occur on the graph path used by the model. |
Python print |
Seeing when tracing occurs. |
tf.print |
Inspecting runtime tensor values at a known point. |
tf.debugging.enable_check_numerics() |
Stopping at an operation that produces a NaN or infinity. |
| TensorBoard Debugger V2 | Investigating unclear origins that need execution history, tensor health, graph, or source context. |
| Input-pipeline analyzer | Checking whether input work is blocking the device. |
| Profiler overview and trace | Examining broader host, device, and timing patterns in slow steps. |
Debugger V2 and Profiler support depends on the installed TensorFlow and TensorBoard releases and the device in use. Check the current documentation and compatibility notes for your environment before relying on a particular API or workflow.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




