Repository navigation
Import extremely slow #368
Description
Activity
Thanks for the detailed log, that helps a lot. The slow line is the one-time Python import of NBI at server startup:
Extension package notebook_intelligence took 80.5213s to importThe underlying reason is that NBI currently imports every provider backend at startup, including
litellm(a very heavy package, roughly 3,400 files), even though you only use Claude. On macOS/Linux that whole import is a few seconds; on Windows a couple of things can blow it up to what you're seeing. Here is how to pin it down and improve it on your side.1. Find the actual bottleneck (2 minutes)
Run this in the same environment and look at the largest entries:
python -X importtime -c "import notebook_intelligence" 2> importtime.log
Then run it a second time (warm vs cold makes a big difference). If one entry near
litellm/get_model_cost_map/httpxdominates, it's a network stall (step 2). If the time is spread across thousands of small submodule imports, it's filesystem/antivirus/OneDrive (steps 3 and 4).2. Stop litellm's network call at import (easy, safe)
litellm fetches a model-cost file over HTTP when it is imported (with a 5s timeout), which can stall badly behind a corporate proxy or TLS inspection. Set this before launching Jupyter:
# Windows (PowerShell) $env:LITELLM_LOCAL_MODEL_COST_MAP = "true"
I measured litellm's import dropping from ~2.0s to ~1.0s with this even on a fast machine with good network, so the win is larger when the network is the problem.
3. Make sure the environment is not under OneDrive
On Windows 11, OneDrive often backs up your Desktop and Documents automatically, and a
.venvliving there gets intercepted on every file read, which is brutal for a 7,000-file import. Check the real path ofmy-jupyterlab: if it resolves under a...\OneDrive\...folder, recreate the environment on a plain local path likeC:\dev\my-jupyterlaband it should speed up significantly.4. If antivirus is the cause, exclude the env folder
Real-time scanning checks each file as it is opened. Adding a Microsoft Defender folder exclusion for your environment (
Windows Security > Virus & threat protection > Manage settings > Exclusions > Add folder) usually helps. One caveat: an excluded folder is no longer scanned, so apply it knowingly to just the env folder. Microsoft's "Dev Drive" performance mode is a safer alternative if you prefer to keep scanning on.On our side
You shouldn't have to import litellm at all when you only use Claude. I opened #370 to make NBI lazy-load provider backends so single-provider users skip that cost entirely. We'll look into it.
One unrelated thing in your log
Your Claude requests are failing auth, separate from the speed issue:
Failed to fetch Claude models: "Could not resolve authentication method..." Claude agent did not reach a terminal connect state within 15.0sThat means Claude isn't authenticated in NBI yet (no Anthropic API key, or the Claude CLI isn't logged in). Worth sorting out separately so Claude mode actually works once startup is faster. The Ollama warning is harmless (it's just that no local Ollama is running).
Hi!
Is this answer AI-generated?
Because I see no reason why my environment should be under OneDrive and I use another antivirus, not Microsoft Defender
Hey @raffaelemancuso, fair question!
Honest answer: yeah, I used Claude to help brainstorm the likely causes, then did my own digging to narrow it down. So the list of suspects was AI-assisted, but here's where I landed: my prime suspect is your antivirus scanning a big pile of small files as they load, not Defender or OneDrive specifically. Let's confirm it rather than assume, though (the importtime step below settles it).
Quick course-corrections based on what you said:
- OneDrive: not under it? Then ignore that one. Just a common Windows gotcha worth checking, not an assumption about your setup.
- Antivirus: doesn't matter that it's not Defender. Pretty much every real-time AV (Norton, McAfee, Kaspersky, Bitdefender, etc.) scans each file as it's opened. NBI currently imports the whole provider stack at startup, and litellm alone ships ~3,400 files, so a lot of small files get opened on import. In bad cases, scanning all of those on open is enough to stretch a few-second import into something much slower. I can't promise it's the entire 80s, which is exactly why the check below matters.
Two things worth trying no matter what:
python -X importtime -c "import notebook_intelligence" 2> importtime.log(run it twice). This basically settles it: if the time is spread across thousands of tiny imports, it's filesystem/AV; if it's concentrated in litellm, it's a network stall. Paste the file here and I'll tell you which.LITELLM_LOCAL_MODEL_COST_MAP=trueto stop litellm's network call at import. I checked, it cuts litellm's load time by about a third on a good connection, and more if the network is the bottleneck.
To confirm the AV theory directly: add an exclusion for your environment folder in whatever AV you use, or briefly turn off real-time protection, and re-time the import. Just scope the exclusion to that folder, flip protection back on after, and if it's a work machine check with IT first.
And honestly, you shouldn't be importing litellm at all when you only use Claude. That's the real fix and it's on us, tracked in #370.
On the Claude auth thing, you're right and I jumped the gun, sorry. That
Failed to fetch Claude models: Could not resolve authentication methodline is a separate, direct Anthropic API call NBI makes just to fill the model dropdown (it needs an Anthropic API key in NBI's settings, or theANTHROPIC_API_KEYenv var). It doesn't mean your CLI login is broken or that Claude mode won't work, those run off your logged-in CLI. Thedid not reach a terminal connect state within 15.0sis a separate agent-startup timeout, not auth. Could be the same slowness, could be the CLI taking too long to come up; if it keeps happening it's worth a closer look.Do I have to delete something before calling
python -X importtime -c "import notebook_intelligence" 2> importtime.logfor a true cold start?Good question, and no, you don't need to delete anything.
Deleting
__pycache__would only reset Python's bytecode compilation, which is a one-time cost after install and doesn't repeat on every launch, so it's very unlikely to be your 80s. Skip it.The thing that actually makes a repeat run faster is the OS file cache (the files are still in RAM from the last run), and partly your AV's scan cache. Those live in memory and clear on a reboot, not by deleting files. (Heads up: some AVs also keep a known-good list on disk that survives reboots, so a reboot resets the OS cache cleanly but only part of the AV side.)
So for the closest thing to a true cold start: reboot, then run the command a few times in a row as the very first thing you do, before opening Jupyter or anything else. The first run is your cold start; the repeats are warm.
How to read it:
- If that first (cold) run is much slower than the warm repeats, the cost is some cold-cache layer (OS file cache, AV scanning, etc.), which fits the "lots of small files getting scanned/read" theory.
- If every run is slow by about the same amount, it's a per-process cost instead, like litellm re-fetching its model-cost map (it does that on every fresh python process, there's no cross-process cache) or just CPU.
To actually pin it on the AV (and fix it at the same time): add an exclusion for your environment folder, or briefly turn off real-time scanning, then re-run. If the cold run speeds up, that's your culprit.
Paste the logs from a couple of runs and I'll tell you which it is.
Thanks, that's exactly what I needed. Two things jump out.
First, the good news: your two runs came in at about 6.5s and 4.25s, not 80s. That points to the original 80s being a one-off cold start, very likely the first import right after install, when bytecode gets compiled across the whole dependency tree, files are read cold off disk, and your AV possibly scans everything freshly written, all of which later runs skip. We never caught an importtime during the actual 80s, so I can't say for certain, but if it ever recurs at anything like that, grab an importtime during the slow start and we'll compare.
Where the time actually goes (warm run, ~4.25s):
- litellm and its submodules: ~1.3s, the single biggest piece, plus openai ~0.4s. You use neither, so that ~1.7s is all overhead you don't need. That's the lazy-load fix I filed in perf(startup): lazy-import provider SDKs to cut server-extension import time #370.
- Your first run is ~2.3s slower than the second, spread across the dependency tree (litellm, pywin32, mcp, anthropic, and so on). That gap is cold-disk reads plus AV possibly scanning of freshly-touched files, which is why a cold start genuinely costs more; an AV exclusion for the env folder should trim it.
- anthropic ~0.5s, plus claude_agent_sdk and mcp, are the parts you actually need.
- About 0.6s is rfc3987_syntax, a grammar-building import that comes in via jsonschema (pulled by jupyter_events and mcp), not litellm. It's a Jupyter-stack cost, so it's not really ours to remove.
You can set
LITELLM_LOCAL_MODEL_COST_MAP=trueto skip a network call litellm makes on every import. It's a small, free win; the fetch is bundled into litellm's own startup time, so you can't read an exact savings off these logs. The bigger win, ~1.7s, is #370 on our side.Short version: the steady-state import is heavy but not crazy (your runs were ~4 to 6.5s), the 80s looks like a cold-start outlier, and the most impactful fix is us not loading litellm for Claude-only users.
- added a commit that references this issue
on Jun 10, 2026 - added a commit that references this issue
on Jun 16, 2026 Any news on this?
Do those commits that defer provider SDK imports to first use work? Because I have the latest version and unused providers are still imported
Following up on the deferred-import question, since it deserved a measurement rather than a guess. You are partly right.
What #370 did fix. On 5.4.x,
import notebook_intelligencepulls in nolitellm, noopenai, and noanthropicSDK. I checkedsys.modulesafter the import in a clean interpreter rather than reading the changelog.What was still loading eagerly: ollama.
AIServiceManagerconstructs all four providers at startup, and the Ollama provider's constructor immediately enumerated local models, which imported theollamapackage and called the Ollama host. That is theFailed to update supported Ollama modelsline in your log, and it ran even with Claude selected and Ollama not installed. Filed as #427, fix in #428.Scale, so I am not overselling it.
import ollamais roughly 0.15 to 0.22s standalone, but only about 35ms marginal once Jupyter has already loadedhttpxandpydantic. That is not your 80 seconds.The part that might actually be your 80 seconds. While bounding those calls I found that the ollama client passes
timeout=Nonestraight to httpx, so it has no timeout at all. If the Ollama host silently drops packets rather than refusing the connection,ollama.list()waits out the operating system's TCP connect timeout. I measured that at 75 seconds on macOS. A refused connection, meaning nothing listening at all, comes back in well under a millisecond, which is the ordinary case and harmless. The expensive case needs something that accepts or blackholes rather than refuses: a firewall that drops, or anOLLAMA_HOSTpointing somewhere unreachable.Before this change, that wait happened during the extension import, which is exactly the line your log attributes the time to.
So, two questions about that machine:
- Is anything listening on port 11434, and if so what?
- Is
OLLAMA_HOSTset in the environment that launches JupyterLab, and does it point at a host or port that might be firewalled or dropping traffic (a VPN-reached host, a container, a stale address)?
I want to be clear that this is a hypothesis and not a diagnosis. I have not reproduced your 80 seconds, and the warm importtime logs you sent (about 4 to 6.5 seconds) contain no stalled connect, so they neither confirm nor rule it out. But if either answer above is yes, then #370 would not have fixed it, lazy loading alone would not fix it either, and the timeout added in #428 would.
If you still have a slow start on the current version, an
-X importtimecaptured during a slow start (not a warm one) would settle it. On my machine the largest remaining entry isrfc3987_syntaxat roughly 0.5s, arriving viajsonschemaandjupyter_events, which is Jupyter stack rather than NBI.This is still an issue. It is loading GitHub Copilot stuff but I only have Claude:
2026-09-19 11:32:42,901 - notebook_intelligence.github_copilot - github_copilot.py - WARNING - Storing the GitHub Copilot token under the default NBI_GH_ACCESS_TOKEN_PASSWORD. Set a per-user password in multi-tenant deployments. [ABRIDGED]\.venv\Lib\site-packages\claude_agent_sdk\types.py:1939: CanUseToolShadowedWarning: can_use_tool will not be invoked for: mcp__nbi__list-available-notebook-kernels, mcp__nbi__create-new-notebook, mcp__nbi__add-markdown-cell, mcp__nbi__add-code-cell, mcp__nbi__get-number-of-cells, mcp__nbi__get-cell-type-and-source, mcp__nbi__get-cell-output, mcp__nbi__set-cell-type-and-source, mcp__nbi__insert-cell, mcp__nbi__save-notebook, mcp__nbi__rename-notebook, mcp__nbi__open-file-in-jupyter-ui. An allowed_tools entry that allows a whole tool auto-approves it before the callback is consulted. To gate every tool call, use a PreToolUse hook; or narrow the entry so calls fall through to can_use_tool. Allow rules from settings files can also shadow the callback but are not visible here. _warn_if_can_use_tool_shadowed(options) 2026-09-19 11:32:43,662 - notebook_intelligence.github_copilot - github_copilot.py - INFO - Using existing GitHub access token 2026-09-19 11:32:43,667 - notebook_intelligence.github_copilot - github_copilot.py - INFO - Refreshing GitHub token@raffaelemancuso Have you pulled the most recent version from main?
@raffaelemancuso Have you pulled the most recent version from main?
yes
Can you try testing what I just pushed to PR #390?

NBI takes 80 seconds to import.
That's extremely slow.
How can I speed it up?
I don't have Github copilot, Amazon BedRock or the like. Only Claude Code.
Here is the full log: