Private credit - #1380
Closed
bryan-xiao97 wants to merge 20 commits into
Closed
Private credit#1380bryan-xiao97 wants to merge 20 commits into
bryan-xiao97 wants to merge 20 commits into
Conversation
resolve_report_filing() in edgar/ai/mcp/tools/selection.py implements the F1 selection rule (accession_number > period > latest, amendments reachable only by accession, never nearest-period fallback) so edgar_fund, edgar_notes and edgar_read can share one filing-selection path instead of re-deriving it. format_source() in base.py builds the rule-3 provenance block, and format_filing_summary() gains an additive period_of_report field. Period matching and listing formatting read report_date (already loaded from the submissions JSON) rather than the period_of_report property, which downloads the entire filing submission per filing; format_source falls back to period_of_report only for the single already-chosen filing, when report_date is unavailable.
Adds edgar/ai/mcp/tools/continuation.py: versioned, self-describing base64url cursors (CursorError with INVALID_CURSOR/CURSOR_MISMATCH/ CURSOR_STALE), record and text pagination helpers, and a thread-safe LRU ResultCache bounded by entry count and bytes. Pure Python, no network or edgar imports beyond base.py's error() helper. This lets a paged MCP call (e.g. 1,481 ARCC holdings) survive a multi-process or restarted MCP server: a cursor carries everything needed to verify it against the current call and re-check its fingerprint against a freshly re-extracted result, rather than depending on this process still having it cached. Out of scope here: wiring this into edgar_fund/edgar_notes/edgar_read (follow-up tasks).
… cap encode_cursor built a cursor from any query without checking MAX_CURSOR_CHARS, so a caller with a large filter could get back a cursor that decode_cursor would immediately reject as INVALID_CURSOR — dead on arrival. It now raises CursorError(INVALID_CURSOR) itself when the encoded result would exceed the 2,048-character cap, so the failure surfaces with a useful message at encode time instead of on the next call.
Add portfolio_investments_from_filing() so callers (and the upcoming MCP
tool) can pull BDC holdings from an explicitly selected filing instead of
always getting the latest one. BDCEntity.portfolio_investments() and
schedule_of_investments() now accept an optional filing= override
(validated against the entity's CIK) and both delegate to the new
function, replacing the duplicated facts-then-statement fallback that used
to live inline. PortfolioInvestments also gains a read-only
extraction_method ('xbrl_facts' | 'statement' | None) so callers can see
which extraction path produced a result.
Ground truth (ARCC 10-Q 0001628280-26-050307) is verified live rather than
via VCR: the full submission is ~88 MB, too large to commit as a cassette.
Added a small-BDC (Princeton Capital Corp) VCR fixture for deterministic
coverage, with its cassette under 10 MB.
The Princeton Capital VCR ground-truth test asserted filing.cik but not filing.accession_number, even though accession is one of the fields the task's ground-truth checklist calls for.
…into edgar_fund bdc_portfolio Fixes the borrower-name bug (checked inv.name instead of company_name, so no borrower was ever emitted), adds accession_number/form/period selection via resolve_report_filing, cursor-based paging via the continuation service, a borrower filter that matches both company_name and identifier, full per-investment fields with honest null-vs-zero serialization, and a paged Schedule-of-Investments text fallback (NO_XBRL when there is no XBRL at all) for filings with no structured holdings.
…ith report-year lookback portfolio_investments_from_filing gains an additive xbrl= parameter so a caller that already parsed a filing's XBRL doesn't force a second parse of the same (potentially huge) submission when extraction yields nothing. lookup_bdc(cik=, ticker=, lookback_years=) checks the latest SEC BDC Report and falls back to prior years: the 2026 report omits Ares Capital Corp (CIK 1287750), present in the 2024 and 2025 reports, so a caller relying only on the latest report (is_bdc_cik, get_bdc_list().get_by_cik/ticker) incorrectly treats an active BDC as unknown. is_bdc_cik is left unchanged.
…, and decompose bdc_portfolio Three fixes to edgar_fund's bdc_portfolio (fix round 1 on Task 4): - BDC resolution silently guessed the wrong entity: fuzzy name search could match an unrelated BDC (e.g. "Ares Capital" matching "Ares Strategic Income Fund"), and the latest-report-only lookup made ARCC unresolvable by ticker, CIK, or accession_number. _find_bdc now resolves a ticker or CIK through lookup_bdc (with report-year lookback), falls back to resolve_company() when a ticker isn't in any report, and only uses fuzzy search for a name -- which must resolve to exactly one candidate or comes back as an AMBIGUOUS_BDC error listing up to 5 candidates, never a silent guess. The accession-only path uses lookup_bdc instead of is_bdc_cik. - The no-structured-data fallback re-parsed a filing's XBRL a second time after extraction already had. The parsed XBRL is now threaded through. - _bdc_portfolio (~210 lines) and _bdc_portfolio_no_structured_data (~107 lines) are decomposed into focused helpers (resolve, decode cursor, load extraction, filter/paginate, build response blocks), including a shared _bdc_identity_fields() that removes a duplicated identity block. No response shape changes except the additive resolved_by key, set when a BDC was resolved via fuzzy search.
Exposes edgar.bdc.nonaccrual.extract_nonaccrual through the MCP, labelling results by evidence_level (investment/aggregate/none) so callers can tell per-investment non-accrual detail apart from a portfolio-level-only value, with fixed interpretation_limits text warning that non-accrual is an accounting status, not a default determination. Reuses the Task 4 selection/pagination/cursor helpers (_resolve_bdc_and_filing, _bdc_identity_fields, _decode_page_cursor/_build_next_cursor/_resolve_cursor_offset, format_source, _cell_number) rather than duplicating them; bdc_portfolio is unchanged. Ground truth: ARCC 10-Q 0001628280-26-050307 -> method footnote, 32 investments, 0 warnings (live network test, no cassette -- submission is ~88 MB). Princeton Capital Corp 10-Q 0001213900-26-090000 genuinely has no extractable non-accrual signal (method none, 2 warnings) -- asserted as measured rather than forced non-empty; VCR cassette is 9.28 MB.
…dgar_read edgar_notes and edgar_read now resolve their filing through the shared resolve_report_filing (accession > period > latest-original), carry a source provenance block, and support cursor continuation: edgar_notes pages a matched note's table rows and context text; edgar_read pages an extracted section's text (cached in text_cache, keyed by accession + section). edgar_read's identifier+form path now excludes amendments, matching resolve_report_filing's latest-original policy. Extracted the fund.py-private cursor try/except wrappers into continuation.py (try_decode_cursor/try_encode_cursor) and added peek_cursor, so edgar_notes can identify which of its two cursor kinds (table vs. context) it was handed before validating it.
…t call _continue_notes built its "expected" tool/query for decode_cursor by peeking the cursor itself, so the topic/detail comparison was checking the cursor against its own stored values -- a cursor minted under one topic (or, for a note's context, one detail level) would silently replay under another instead of failing CURSOR_MISMATCH. Thread the current call's topic/detail into _continue_notes and build the expected query from those plus the cursor's own positional note/table indices, so a topic or detail change between calls is now rejected for real.
Adds a "Credit evidence (BDC loans)" workflow to SERVER_INSTRUCTIONS, the fund_analysis prompt, and docs/ai/mcp-tools.md/mcp-workflows.md covering edgar_fund's bdc_portfolio/bdc_nonaccrual (accession_number, form, period, borrower, cursor, include_untyped) and edgar_notes/edgar_read's period and cursor continuation, plus the filing-selection and cursor error codes. Adds a filing-scoped Python usage pattern to the funds skill.yaml and a fast test (tests/test_mcp_docs_examples.py) that checks every documented example's parameter names and enum values against each tool's registered MCP schema. Two changelog fragments record: filing-scoped BDC loan/non-accrual evidence with continuation (ARCC's 2026-06-30 10-Q: 1,481 holdings, all now reachable), and the earlier bdc_portfolio fix where every returned investment's borrower name was null because the record checked a `name` attribute instead of `company_name`.
…phase 2) Adds edgar_document (list/search/read) so agents can find, search and read a specific exhibit with document identity preserved: SEC Archives-only URL input, ambiguous exhibit types rejected with candidates, 6,000-character pages with cursor continuation, and char_offset locators from Filing.grep. edgar_text_search now surfaces the matched file's type, description and document id; edgar_filing keeps the document named in a URL and reports the period and filing URL. Committed as-is for review; a QA sweep found open issues in this change that follow-up commits address.
Fixes the CI-failing raw-ValueError ratchet (3 new sites converted to ValidationError) plus five QA findings from qa-phase1-report.md: - P1-M1: resolve_report_filing's latest path retries with a full history load when the recent page has no match, mirroring the period path and Entity.latest(). - P1-M3: a malformed accession_number (a CIK, a name, stray whitespace, the 18-digit form) is validated by shape before ever reaching edgar.find(), and a find() result that isn't a Filing is FILING_NOT_FOUND instead of crashing downstream. - P1-L2: an explicit amendment form on the period/latest paths is now INVALID_ARGUMENTS instead of silently returning the original filing. - P1-L1: get_latest_bdc_report_year is lru_cached, and lookup_bdc's latest-year lookup shares get_bdc_list()'s cache entry instead of fetching the same CSV under a second key. - P1-M5 (library part): lookup_bdc matches now carry an additive report_year attribute. Also adds peek_cursor_accession (continuation.py) for the cursor-only continuation rule that Q2/Q3 wire into edgar_fund/edgar_notes/edgar_read, and ValidationError guards on paginate/paginate_text for the deferred negative-offset/zero-budget minors. fund.py, notes.py, reader.py and document.py wiring are out of scope here (Tasks Q2-Q4).
…l totals (QA fix wave Q2)
- P1-H1: peek the cursor's tool first; edgar_fund:bdc_soi_text cursors go
straight to the SOI text fallback, which checks text_cache before parsing
XBRL. A no-structured-holdings extraction outcome is cached too. text_page
no longer exposes next_offset (P1-L6).
- P1-H2 / rule 7b: a cursor alone continues bdc_portfolio (including the SOI
fallback) and bdc_nonaccrual. The filing comes from the cursor accession;
borrower/include_untyped default to the cursor query; a differing value or
accession_number is CURSOR_MISMATCH; an identifier for another CIK is
SELECTION_MISMATCH; form/period are ignored with a cursor.
- P1-M2: short names fall back to name search after a ticker miss; a search
resolves only on one exact name match or one hit >= 95 with the runner-up
< 90, else AMBIGUOUS_BDC (up to 5). find_bdc/BDCSearchIndex gain an
additive lookback_years so "Ares Capital Corp" finds ARCC (CIK 1287750,
only in the 2025 report); one row per CIK, latest year wins.
- P1-M4: num_nonaccrual and nonaccrual page totals are null without
investment-level evidence; total_fair_value/total_cost/filtered_totals sum
known values only (null when none); SOI fallback total_investments is null.
- P1-M5: a lookback-year BDC match derives is_active from the company's own
latest filing (18-month rule, filed_within_active_window), null if
unavailable; responses add bdc_report_year.
- P1-L10: include_untyped is parsed strictly ("true"/"false" strings or
bools, else INVALID_ARGUMENTS); a numeric identifier that is a real
non-BDC company is NOT_A_BDC (an unknown CIK stays COMPANY_NOT_FOUND).
…dc (Q2 fix round 1) fund.py was 1,737 lines. It now keeps the @tool registration, argument dispatch and the non-BDC fund actions (550 lines); the BDC code moves verbatim into a private package: - bdc/identity.py (424): bdc_search action, identifier resolution, is_active, identity fields, filing selection - bdc/paging.py (137): cursor routing and page helpers shared by both actions - bdc/portfolio.py (464): bdc_portfolio and its SOI text fallback - bdc/nonaccrual.py (238): bdc_nonaccrual No behaviour change: every moved definition is AST-identical to the previous one apart from renames (names used across modules drop their leading underscore) and one docstring; edgar_fund's schema and description are byte-identical. Tests now patch names where they are looked up (bdc.identity.resolve_bdc / resolve_company), the patch helpers record calls, and new TestPatchSeamsAreLive tests assert the resolution, selection, extraction, resolve_company and lookup_bdc seams are actually called.
# Conflicts: # CHANGELOG.md # docs/common-pitfalls.md # docs/configuration.md # docs/guides/bdc-guide.md # docs/guides/thirteenf-data-object-guide.md # docs/guides/track-form4.md # docs/installation.md # docs/quickstart.md # docs/resources/performance.md # docs/resources/sec-compliance.md # docs/upgrade/6.0.md # edgar/bdc/investments.py # examples/scripts/advanced/enterprise_config.py
dgunning
requested changes
Sep 29, 2026
dgunning
left a comment
Owner
There was a problem hiding this comment.
Thanks @bryan-xiao97. There's a lot of solid work in here: filing selection, cursor continuation, BDC nonaccrual and portfolio paging, and the edgar_document tool. I'd like to get it in, but I can't review or merge the PR in its current shape. It touches 962 files (+510k/−3.2M lines), and most of that has nothing to do with the feature.
What's blocking
- The cleanup commits.
57185694("cleaned up non-essential files") deletes 909 tracked files:docs/,scripts/and the fixtures underdata/. 111 of thosedata/files are loaded by the test suite or the library, so tests would break. It also rewritesREADME.mdand renamesCLAUDE.md→AGENTS.md. Please drop57185694anda1bff7b5entirely. None of those files should change in this work. - Internal planning docs.
.agents-context/(specs, plans, HTML design docs) is committed. Adding it to.gitignoreafterwards doesn't remove the files already committed, so please take them out of the branch. - Conflicts with main. The branch starts from 2026-09-08, and
edgar/bdc/investments.pyhas changed since (#1374, #1375). Please mergemaininto your branch rather than rebasing it. - Size. Four cassettes are about 99k lines each. Please record only the requests each test needs, or use a small offline fixture.
How to split it
It'll be much quicker to review as a few focused PRs, each with a description:
Filing.grep()changes (render_markdown,max_matches,max_match_chars,regex_timeout) and the newregexdependency. This touches the core API, so it should go in its own PR.- The shared selection and continuation services (
selection.py,continuation.py) - The
edgar_fundBDC actions (tools/bdc/, plus theedgar/bdcchanges) edgar_document
If you could open an issue or discussion outlining the private-credit MCP design first, we can agree on the tool surface before you split the code. Happy to help work out where the seams go.
This was referenced Oct 1, 2026
dgunning
added a commit
that referenced
this pull request
Oct 1, 2026
…t row expires (huul) (#1396) Since #1146 the register unions the last three SEC BDC reports and keeps each registrant's newest row. 22 registrants the 2026 report dropped (Ares Capital, Gladstone Capital, Great Elm, Goldman Sachs Private Credit, ...) keep 2025 rows dated ~2025-05-29, which leave the 18-month window around 2026-11-29: is_active would have flipped to False for BDCs that file every quarter. Rows now carry in_latest_report. is_active: - report row inside the window -> True, no request; - row from the newest report, outside the window -> False (the SEC's own current snapshot; these are the 1995-dated dormant registrants); - row from an older report, outside the window -> the company's newest filing date from its submissions JSON, lru-cached per CIK, so at most 22 requests per session today rather than one per stale row (52). An unreachable SEC falls back to the report row with a warning; any other error propagates rather than reading as "inactive". An explicit get_bdc_list(year=...) older than the newest report marks every row not-latest, since it is a snapshot of then. PR #1380's version exposed the window rule as a helper for callers but left is_active on the stale row. Tests freeze the module clock at 2026-12-01: offline cases for each branch, the register flag, the submissions parse, and a live ARCC/MAIN check (ARCC's own latest filing >= its 2026-07-29 10-Q). Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
dgunning
added a commit
that referenced
this pull request
Oct 1, 2026
…s keep source offsets (5qqy, 6th2) (#1397) 5qqy: Filing._attachment_matches tested document.upper() in doc_type, a substring match, so "EX-10.1" also selected EX-10.10..EX-10.19. Replaced by _select_attachments, which sees the whole list: an exact filename wins, then an exact document type, and the old substring match applies only when nothing matched exactly, so "EX-10" still selects every EX-10 exhibit. PR #1380 added only an exact-filename check ahead of the substring test, which leaves the type collision in place. SeeQC's S-1 (0001213900-26-073222): "agreement" in EX-10.1 is 133 matches; the old filter returned 378, adding EX-10.10/11/12. 6th2: the literal branch of _grep_text found positions in text.lower() and sliced the original text with them. "İ".lower() is two characters, so each one before a match shifted it by one. Literal search now runs re.escape(pattern) with IGNORECASE over the original text. It keeps the old scan's overlapping matches ("aa" in "aaa" is two) by searching again from m.start() + 1 rather than using finditer; regex mode is unchanged. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
dgunning
added a commit
that referenced
this pull request
Oct 1, 2026
…n BDCs (1c6z) (#1398) _bdc_portfolio fell back to find_bdc(identifier, top_n=1)[0] when the identifier was neither a ticker nor a CIK. find_bdc scores every whole-word hit near 100 (it only nudges toward the closest full name), so the top hit identified nothing: "Golub" matches four Golub BDCs at 99.4-100 and silently became GOLUB CAPITAL BDC. _pick_bdc_search_winner resolves a name only when (a) exactly one result equals the query ignoring case, punctuation and a trailing corporate suffix, or (b) the top hit scores >= 95 and the runner-up < 90. Otherwise the tool returns AMBIGUOUS_BDC with up to five "name: identifier='CIK'" suggestions (CIK, not ticker: non-traded BDCs have none). The rule is PR #1380's (d12c057) with one change: it compared full names literally, which would have made "Ares Capital" and "Prospect Capital" ambiguous (siblings score 99.3; neither equals "...CORP"), though both resolve correctly today. Ignoring the suffix keeps them resolving. Measured on 17 queries: Golub, Ares, Blackstone, Blue Owl, Gladstone, Oaktree, Monroe, Barings, Hercules -> ambiguous; Ares Capital, Main Street, Prospect Capital, Hercules Capital, WhiteHorse, Saratoga -> one. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.