<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>MOONGCHI — English Articles</title>
    <description>Practical notes on software engineering, data systems, and machine learning.</description>
    <language>en</language>
    <link>https://berrrrr.github.io/en/</link>
    <atom:link href="https://berrrrr.github.io/en/feed.xml" rel="self" type="application/rss+xml"/>
    <pubDate>Thu, 01 Oct 2026 12:32:04 +0000</pubDate>
    <lastBuildDate>Thu, 01 Oct 2026 12:32:04 +0000</lastBuildDate>
    
    
      <item>
        <title>ONNX Runtime Troubleshooting: CUDA, cuDNN, Memory, and Inference Performance</title>
        <description>&lt;blockquote&gt;
  &lt;p&gt;This article is part of my &lt;strong&gt;ONNX series&lt;/strong&gt;. It collects practical lessons from troubleshooting ONNX Runtime in GPU serving environments.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3 id=&quot;1-libcudnnso8-cannot-open-shared-object-file&quot;&gt;1. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;libcudnn.so.8: cannot open shared object file&lt;/code&gt;&lt;/h3&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;[E:onnxruntime:Default, provider_bridge_ort.cc:1745 TryGetProviderInfo_CUDA]
/onnxruntime_src/onnxruntime/core/session/provider_bridge_ort.cc:1426
onnxruntime::Provider&amp;amp; onnxruntime::ProviderLibrary::Get()
[ONNXRuntimeError] : 1 : FAIL : Failed to load library
libonnxruntime_providers_cuda.so with error: libcudnn.so.8:
cannot open shared object file: No such file or directory
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;First, add the CUDA library directory to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LD_LIBRARY_PATH&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;/usr/local/cuda/lib64:&lt;span class=&quot;nv&quot;&gt;$LD_LIBRARY_PATH&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If the error remains, check that your CUDA and cuDNN versions match the versions supported by ONNX Runtime. The &lt;a href=&quot;https://onnxruntime.ai/docs/execution-providers/CUDA-ExecutionProvider.html#requirements&quot;&gt;CUDA Execution Provider requirements&lt;/a&gt; contain the compatibility table.&lt;/p&gt;

&lt;h3 id=&quot;2-cuda-out-of-memory&quot;&gt;2. CUDA out of memory&lt;/h3&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 576.00 MiB. GPU
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;I kept encountering this OOM error. Oddly, after adding per-task memory logging with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pynvml&lt;/code&gt;, the error stopped occurring. I could not establish a clear causal relationship, so this observation needs further verification rather than being treated as a fix.&lt;/p&gt;

&lt;h3 id=&quot;3-non-zero-status-code-returned-while-running&quot;&gt;3. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Non-zero status code returned while running&lt;/code&gt;&lt;/h3&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;[E:onnxruntime:, sequential_executor.cc:514 ExecuteKernel]
Non-zero status code returned while running Softmax node.
Name:&apos;/encoder/blocks/blocks.0/attn/Softmax&apos;
Status Message: /onnxruntime_src/onnxruntime/core/framework/bfc_arena.cc:376
Failed to allocate memory for requested buffer of size 1775665152

onnxruntime.capi.onnxruntime_pybind11_state.RuntimeException:
[ONNXRuntimeError] : 6 : RUNTIME_EXCEPTION : Non-zero status code returned
while running MatMul node. Name:&apos;/encoder/blocks/blocks.0/attn/MatMul&apos;
Status Message: /onnxruntime_src/onnxruntime/core/framework/bfc_arena.cc:376
Failed to allocate memory for requested buffer of size 2500263936
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;In my case, this occurred in a PARSeq text recognition model. I had configured dynamic axes and enabled batch inference, expecting it to improve throughput. Switching from batch inference to single-item inference resolved the allocation failure.&lt;/p&gt;

&lt;h3 id=&quot;4-libcudnn_advso9-cannot-open-shared-object-file&quot;&gt;4. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;libcudnn_adv.so.9: cannot open shared object file&lt;/code&gt;&lt;/h3&gt;

&lt;p&gt;The directory containing the cuDNN libraries installed in the virtual environment must be included in the library path. Depending on the environment, it may be under one of these locations:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;/home/user/.local/lib/python3.11/
/home/user/miniconda3/envs/{env_name}/lib/python3.11/
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;For a Poetry-created virtual environment in my container, I added this path:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;:/usr/src/app/.venv/lib/python3.12/site-packages/nvidia/cudnn/lib&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;5-libnvrtcso12-cannot-open-shared-object-file&quot;&gt;5. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;libnvrtc.so.12: cannot open shared object file&lt;/code&gt;&lt;/h3&gt;

&lt;p&gt;In my environment, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;libnvrtc.so.12&lt;/code&gt; was located here:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;/usr/src/app/.venv/lib/python3.12/site-packages/nvidia/cuda_nvrtc/lib
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;I added that directory as well:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;:/usr/src/app/.venv/lib/python3.12/site-packages/nvidia/cuda_nvrtc/lib&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;You can diagnose the problem in the following order.&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;
    &lt;p&gt;Check the CUDA Toolkit version:&lt;/p&gt;

    &lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;nvcc &lt;span class=&quot;nt&quot;&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;    &lt;/div&gt;

    &lt;p&gt;If CUDA 12.x is installed, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;libnvrtc.so.12&lt;/code&gt; should be available.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Find the library:&lt;/p&gt;

    &lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;locate libnvrtc.so
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;    &lt;/div&gt;

    &lt;p&gt;Or:&lt;/p&gt;

    &lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;find /usr &lt;span class=&quot;nt&quot;&gt;-name&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;libnvrtc.so*&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;    &lt;/div&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Add its directory to the environment. For example:&lt;/p&gt;

    &lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;/usr/local/cuda-12.x/lib64:&lt;span class=&quot;nv&quot;&gt;$LD_LIBRARY_PATH&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;    &lt;/div&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;If the file is missing on Ubuntu, install the appropriate CUDA package for your environment. For example:&lt;/p&gt;

    &lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;apt-get &lt;span class=&quot;nb&quot;&gt;install &lt;/span&gt;nvidia-cuda-toolkit
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;    &lt;/div&gt;

    &lt;p&gt;Or install the matching CUDA 12.x runtime library package:&lt;/p&gt;

    &lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;apt-get &lt;span class=&quot;nb&quot;&gt;install &lt;/span&gt;cuda-libraries-12-&lt;span class=&quot;k&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;nt&quot;&gt;-nvrtc&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;    &lt;/div&gt;
  &lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Package names and installation methods vary by CUDA repository and Ubuntu version, so confirm them against the NVIDIA installation documentation for your system.&lt;/p&gt;

&lt;h3 id=&quot;model-serving-lessons&quot;&gt;Model-serving lessons&lt;/h3&gt;

&lt;h4 id=&quot;differences-in-inference-accuracy&quot;&gt;Differences in inference accuracy&lt;/h4&gt;

&lt;p&gt;The accuracy difference ultimately came from the CUDA Toolkit version, cuDNN version, and GPU model, even though the other library versions were identical.&lt;/p&gt;

&lt;p&gt;The training and serving environments used the same GPU base image, so I initially assumed their CUDA stacks were identical. However, extra CUDA Toolkit and cuDNN packages had been installed during training. Check the versions that PyTorch actually uses:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;torch&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;cuda toolkit:&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;torch&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;version&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cuda&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;cudnn:&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;torch&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;backends&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cudnn&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;version&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Matching these versions is important. Even after aligning the software environment, different physical GPUs—for example, an NVIDIA L4 and an A10G—can still produce small numerical differences.&lt;/p&gt;

&lt;h4 id=&quot;differences-in-inference-speed&quot;&gt;Differences in inference speed&lt;/h4&gt;

&lt;p&gt;First, ONNX Runtime could not load the required cuDNN libraries and therefore did not reach its expected performance. The relevant errors were:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;libcudnn_adv.so.9: cannot open shared object file: No such file or directory
libnvrtc.so.12: cannot open shared object file: No such file or directory
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The directories containing the installed CUDA libraries had to be added to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LD_LIBRARY_PATH&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;:/usr/src/app/.venv/lib/python3.12/site-packages/nvidia/cudnn/lib&quot;&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LD_LIBRARY_PATH&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;:/usr/src/app/.venv/lib/python3.12/site-packages/nvidia/cuda_nvrtc/lib&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;These paths belong to a Poetry-created virtual environment in this particular container. Locate the libraries and use the paths that match your own environment.&lt;/p&gt;

&lt;p&gt;The second bottleneck was CPU capacity. Even with all library versions aligned, inference was much slower in the serving environment. Parts of an ONNX graph can execute on the CPU, so the number of CPU cores still matters.&lt;/p&gt;

&lt;p&gt;The training environment exposed 190 CPU cores and completed inference in roughly 0.02 seconds. The serving environment exposed only four cores and took about 0.2 seconds. Increasing the serving environment to 48 cores brought inference time close to the training environment.&lt;/p&gt;
</description>
        <pubDate>Fri, 19 Sep 2025 00:00:00 +0000</pubDate>
        <link>https://berrrrr.github.io/en/onnx-runtime-troubleshooting-lessons/</link>
        <guid isPermaLink="true">https://berrrrr.github.io/en/onnx-runtime-troubleshooting-lessons/</guid>
        <category>mlops</category>
        <category>programming</category>
      </item>
    
  </channel>
</rss>
