Skip to content

Private credit - #1380

Closed
bryan-xiao97 wants to merge 20 commits into
dgunning:mainfrom
bryan-xiao97:private-credit
Closed

bryan-xiao97 wants to merge 20 commits into
dgunning:mainfrom
bryan-xiao97:private-credit

Conversation

@bryan-xiao97

Copy link
Copy Markdown

No description provided.

resolve_report_filing() in edgar/ai/mcp/tools/selection.py implements the F1
selection rule (accession_number > period > latest, amendments reachable
only by accession, never nearest-period fallback) so edgar_fund, edgar_notes
and edgar_read can share one filing-selection path instead of re-deriving it.
format_source() in base.py builds the rule-3 provenance block, and
format_filing_summary() gains an additive period_of_report field.

Period matching and listing formatting read report_date (already loaded from
the submissions JSON) rather than the period_of_report property, which
downloads the entire filing submission per filing; format_source falls back
to period_of_report only for the single already-chosen filing, when
report_date is unavailable.
Adds edgar/ai/mcp/tools/continuation.py: versioned, self-describing
base64url cursors (CursorError with INVALID_CURSOR/CURSOR_MISMATCH/
CURSOR_STALE), record and text pagination helpers, and a thread-safe LRU
ResultCache bounded by entry count and bytes. Pure Python, no network or
edgar imports beyond base.py's error() helper.

This lets a paged MCP call (e.g. 1,481 ARCC holdings) survive a
multi-process or restarted MCP server: a cursor carries everything needed
to verify it against the current call and re-check its fingerprint against
a freshly re-extracted result, rather than depending on this process still
having it cached.

Out of scope here: wiring this into edgar_fund/edgar_notes/edgar_read
(follow-up tasks).
… cap

encode_cursor built a cursor from any query without checking
MAX_CURSOR_CHARS, so a caller with a large filter could get back a cursor
that decode_cursor would immediately reject as INVALID_CURSOR — dead on
arrival. It now raises CursorError(INVALID_CURSOR) itself when the encoded
result would exceed the 2,048-character cap, so the failure surfaces with a
useful message at encode time instead of on the next call.
Add portfolio_investments_from_filing() so callers (and the upcoming MCP
tool) can pull BDC holdings from an explicitly selected filing instead of
always getting the latest one. BDCEntity.portfolio_investments() and
schedule_of_investments() now accept an optional filing= override
(validated against the entity's CIK) and both delegate to the new
function, replacing the duplicated facts-then-statement fallback that used
to live inline. PortfolioInvestments also gains a read-only
extraction_method ('xbrl_facts' | 'statement' | None) so callers can see
which extraction path produced a result.

Ground truth (ARCC 10-Q 0001628280-26-050307) is verified live rather than
via VCR: the full submission is ~88 MB, too large to commit as a cassette.
Added a small-BDC (Princeton Capital Corp) VCR fixture for deterministic
coverage, with its cassette under 10 MB.
The Princeton Capital VCR ground-truth test asserted filing.cik but not
filing.accession_number, even though accession is one of the fields the
task's ground-truth checklist calls for.
…into edgar_fund bdc_portfolio

Fixes the borrower-name bug (checked inv.name instead of company_name, so
no borrower was ever emitted), adds accession_number/form/period selection
via resolve_report_filing, cursor-based paging via the continuation
service, a borrower filter that matches both company_name and identifier,
full per-investment fields with honest null-vs-zero serialization, and a
paged Schedule-of-Investments text fallback (NO_XBRL when there is no XBRL
at all) for filings with no structured holdings.
…ith report-year lookback

portfolio_investments_from_filing gains an additive xbrl= parameter so a
caller that already parsed a filing's XBRL doesn't force a second parse of
the same (potentially huge) submission when extraction yields nothing.

lookup_bdc(cik=, ticker=, lookback_years=) checks the latest SEC BDC Report
and falls back to prior years: the 2026 report omits Ares Capital Corp
(CIK 1287750), present in the 2024 and 2025 reports, so a caller relying
only on the latest report (is_bdc_cik, get_bdc_list().get_by_cik/ticker)
incorrectly treats an active BDC as unknown. is_bdc_cik is left unchanged.
…, and decompose bdc_portfolio

Three fixes to edgar_fund's bdc_portfolio (fix round 1 on Task 4):

- BDC resolution silently guessed the wrong entity: fuzzy name search could
  match an unrelated BDC (e.g. "Ares Capital" matching "Ares Strategic
  Income Fund"), and the latest-report-only lookup made ARCC unresolvable
  by ticker, CIK, or accession_number. _find_bdc now resolves a ticker or
  CIK through lookup_bdc (with report-year lookback), falls back to
  resolve_company() when a ticker isn't in any report, and only uses fuzzy
  search for a name -- which must resolve to exactly one candidate or comes
  back as an AMBIGUOUS_BDC error listing up to 5 candidates, never a silent
  guess. The accession-only path uses lookup_bdc instead of is_bdc_cik.
- The no-structured-data fallback re-parsed a filing's XBRL a second time
  after extraction already had. The parsed XBRL is now threaded through.
- _bdc_portfolio (~210 lines) and _bdc_portfolio_no_structured_data
  (~107 lines) are decomposed into focused helpers (resolve, decode cursor,
  load extraction, filter/paginate, build response blocks), including a
  shared _bdc_identity_fields() that removes a duplicated identity block.

No response shape changes except the additive resolved_by key, set when a
BDC was resolved via fuzzy search.
Exposes edgar.bdc.nonaccrual.extract_nonaccrual through the MCP, labelling
results by evidence_level (investment/aggregate/none) so callers can tell
per-investment non-accrual detail apart from a portfolio-level-only value,
with fixed interpretation_limits text warning that non-accrual is an
accounting status, not a default determination.

Reuses the Task 4 selection/pagination/cursor helpers (_resolve_bdc_and_filing,
_bdc_identity_fields, _decode_page_cursor/_build_next_cursor/_resolve_cursor_offset,
format_source, _cell_number) rather than duplicating them; bdc_portfolio is
unchanged.

Ground truth: ARCC 10-Q 0001628280-26-050307 -> method footnote, 32
investments, 0 warnings (live network test, no cassette -- submission is
~88 MB). Princeton Capital Corp 10-Q 0001213900-26-090000 genuinely has no
extractable non-accrual signal (method none, 2 warnings) -- asserted as
measured rather than forced non-empty; VCR cassette is 9.28 MB.
…dgar_read

edgar_notes and edgar_read now resolve their filing through the shared
resolve_report_filing (accession > period > latest-original), carry a
source provenance block, and support cursor continuation: edgar_notes
pages a matched note's table rows and context text; edgar_read pages an
extracted section's text (cached in text_cache, keyed by accession +
section). edgar_read's identifier+form path now excludes amendments,
matching resolve_report_filing's latest-original policy.

Extracted the fund.py-private cursor try/except wrappers into
continuation.py (try_decode_cursor/try_encode_cursor) and added
peek_cursor, so edgar_notes can identify which of its two cursor kinds
(table vs. context) it was handed before validating it.
…t call

_continue_notes built its "expected" tool/query for decode_cursor by
peeking the cursor itself, so the topic/detail comparison was checking
the cursor against its own stored values -- a cursor minted under one
topic (or, for a note's context, one detail level) would silently
replay under another instead of failing CURSOR_MISMATCH. Thread the
current call's topic/detail into _continue_notes and build the expected
query from those plus the cursor's own positional note/table indices,
so a topic or detail change between calls is now rejected for real.
Adds a "Credit evidence (BDC loans)" workflow to SERVER_INSTRUCTIONS, the
fund_analysis prompt, and docs/ai/mcp-tools.md/mcp-workflows.md covering
edgar_fund's bdc_portfolio/bdc_nonaccrual (accession_number, form, period,
borrower, cursor, include_untyped) and edgar_notes/edgar_read's period and
cursor continuation, plus the filing-selection and cursor error codes.
Adds a filing-scoped Python usage pattern to the funds skill.yaml and a
fast test (tests/test_mcp_docs_examples.py) that checks every documented
example's parameter names and enum values against each tool's registered
MCP schema.

Two changelog fragments record: filing-scoped BDC loan/non-accrual
evidence with continuation (ARCC's 2026-06-30 10-Q: 1,481 holdings, all
now reachable), and the earlier bdc_portfolio fix where every returned
investment's borrower name was null because the record checked a `name`
attribute instead of `company_name`.
…phase 2)

Adds edgar_document (list/search/read) so agents can find, search and read a
specific exhibit with document identity preserved: SEC Archives-only URL
input, ambiguous exhibit types rejected with candidates, 6,000-character pages
with cursor continuation, and char_offset locators from Filing.grep.
edgar_text_search now surfaces the matched file's type, description and
document id; edgar_filing keeps the document named in a URL and reports the
period and filing URL.

Committed as-is for review; a QA sweep found open issues in this change that
follow-up commits address.
Fixes the CI-failing raw-ValueError ratchet (3 new sites converted to
ValidationError) plus five QA findings from qa-phase1-report.md:

- P1-M1: resolve_report_filing's latest path retries with a full history
  load when the recent page has no match, mirroring the period path and
  Entity.latest().
- P1-M3: a malformed accession_number (a CIK, a name, stray whitespace, the
  18-digit form) is validated by shape before ever reaching edgar.find(),
  and a find() result that isn't a Filing is FILING_NOT_FOUND instead of
  crashing downstream.
- P1-L2: an explicit amendment form on the period/latest paths is now
  INVALID_ARGUMENTS instead of silently returning the original filing.
- P1-L1: get_latest_bdc_report_year is lru_cached, and lookup_bdc's
  latest-year lookup shares get_bdc_list()'s cache entry instead of
  fetching the same CSV under a second key.
- P1-M5 (library part): lookup_bdc matches now carry an additive
  report_year attribute.

Also adds peek_cursor_accession (continuation.py) for the cursor-only
continuation rule that Q2/Q3 wire into edgar_fund/edgar_notes/edgar_read,
and ValidationError guards on paginate/paginate_text for the deferred
negative-offset/zero-budget minors.

fund.py, notes.py, reader.py and document.py wiring are out of scope here
(Tasks Q2-Q4).
…l totals (QA fix wave Q2)

- P1-H1: peek the cursor's tool first; edgar_fund:bdc_soi_text cursors go
  straight to the SOI text fallback, which checks text_cache before parsing
  XBRL. A no-structured-holdings extraction outcome is cached too. text_page
  no longer exposes next_offset (P1-L6).
- P1-H2 / rule 7b: a cursor alone continues bdc_portfolio (including the SOI
  fallback) and bdc_nonaccrual. The filing comes from the cursor accession;
  borrower/include_untyped default to the cursor query; a differing value or
  accession_number is CURSOR_MISMATCH; an identifier for another CIK is
  SELECTION_MISMATCH; form/period are ignored with a cursor.
- P1-M2: short names fall back to name search after a ticker miss; a search
  resolves only on one exact name match or one hit >= 95 with the runner-up
  < 90, else AMBIGUOUS_BDC (up to 5). find_bdc/BDCSearchIndex gain an
  additive lookback_years so "Ares Capital Corp" finds ARCC (CIK 1287750,
  only in the 2025 report); one row per CIK, latest year wins.
- P1-M4: num_nonaccrual and nonaccrual page totals are null without
  investment-level evidence; total_fair_value/total_cost/filtered_totals sum
  known values only (null when none); SOI fallback total_investments is null.
- P1-M5: a lookback-year BDC match derives is_active from the company's own
  latest filing (18-month rule, filed_within_active_window), null if
  unavailable; responses add bdc_report_year.
- P1-L10: include_untyped is parsed strictly ("true"/"false" strings or
  bools, else INVALID_ARGUMENTS); a numeric identifier that is a real
  non-BDC company is NOT_A_BDC (an unknown CIK stays COMPANY_NOT_FOUND).
…dc (Q2 fix round 1)

fund.py was 1,737 lines. It now keeps the @tool registration, argument
dispatch and the non-BDC fund actions (550 lines); the BDC code moves
verbatim into a private package:

- bdc/identity.py (424): bdc_search action, identifier resolution,
  is_active, identity fields, filing selection
- bdc/paging.py (137): cursor routing and page helpers shared by both actions
- bdc/portfolio.py (464): bdc_portfolio and its SOI text fallback
- bdc/nonaccrual.py (238): bdc_nonaccrual

No behaviour change: every moved definition is AST-identical to the
previous one apart from renames (names used across modules drop their
leading underscore) and one docstring; edgar_fund's schema and
description are byte-identical. Tests now patch names where they are
looked up (bdc.identity.resolve_bdc / resolve_company), the patch
helpers record calls, and new TestPatchSeamsAreLive tests assert the
resolution, selection, extraction, resolve_company and lookup_bdc
seams are actually called.
# Conflicts:
#	CHANGELOG.md
#	docs/common-pitfalls.md
#	docs/configuration.md
#	docs/guides/bdc-guide.md
#	docs/guides/thirteenf-data-object-guide.md
#	docs/guides/track-form4.md
#	docs/installation.md
#	docs/quickstart.md
#	docs/resources/performance.md
#	docs/resources/sec-compliance.md
#	docs/upgrade/6.0.md
#	edgar/bdc/investments.py
#	examples/scripts/advanced/enterprise_config.py

@dgunning dgunning left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @bryan-xiao97. There's a lot of solid work in here: filing selection, cursor continuation, BDC nonaccrual and portfolio paging, and the edgar_document tool. I'd like to get it in, but I can't review or merge the PR in its current shape. It touches 962 files (+510k/−3.2M lines), and most of that has nothing to do with the feature.

What's blocking

  1. The cleanup commits. 57185694 ("cleaned up non-essential files") deletes 909 tracked files: docs/, scripts/ and the fixtures under data/. 111 of those data/ files are loaded by the test suite or the library, so tests would break. It also rewrites README.md and renames CLAUDE.md → AGENTS.md. Please drop 57185694 and a1bff7b5 entirely. None of those files should change in this work.
  2. Internal planning docs. .agents-context/ (specs, plans, HTML design docs) is committed. Adding it to .gitignore afterwards doesn't remove the files already committed, so please take them out of the branch.
  3. Conflicts with main. The branch starts from 2026-09-08, and edgar/bdc/investments.py has changed since (#1374, #1375). Please merge main into your branch rather than rebasing it.
  4. Size. Four cassettes are about 99k lines each. Please record only the requests each test needs, or use a small offline fixture.

How to split it

It'll be much quicker to review as a few focused PRs, each with a description:

  • Filing.grep() changes (render_markdown, max_matches, max_match_chars, regex_timeout) and the new regex dependency. This touches the core API, so it should go in its own PR.
  • The shared selection and continuation services (selection.py, continuation.py)
  • The edgar_fund BDC actions (tools/bdc/, plus the edgar/bdc changes)
  • edgar_document

If you could open an issue or discussion outlining the private-credit MCP design first, we can agree on the tool surface before you split the code. Happy to help work out where the seams go.

dgunning added a commit that referenced this pull request Oct 1, 2026
…t row expires (huul) (#1396)

Since #1146 the register unions the last three SEC BDC reports and keeps
each registrant's newest row. 22 registrants the 2026 report dropped
(Ares Capital, Gladstone Capital, Great Elm, Goldman Sachs Private Credit,
...) keep 2025 rows dated ~2025-05-29, which leave the 18-month window
around 2026-11-29: is_active would have flipped to False for BDCs that
file every quarter.

Rows now carry in_latest_report. is_active:
- report row inside the window -> True, no request;
- row from the newest report, outside the window -> False (the SEC's own
  current snapshot; these are the 1995-dated dormant registrants);
- row from an older report, outside the window -> the company's newest
  filing date from its submissions JSON, lru-cached per CIK, so at most
  22 requests per session today rather than one per stale row (52).
An unreachable SEC falls back to the report row with a warning; any other
error propagates rather than reading as "inactive". An explicit
get_bdc_list(year=...) older than the newest report marks every row
not-latest, since it is a snapshot of then.

PR #1380's version exposed the window rule as a helper for callers but
left is_active on the stale row.

Tests freeze the module clock at 2026-12-01: offline cases for each
branch, the register flag, the submissions parse, and a live ARCC/MAIN
check (ARCC's own latest filing >= its 2026-07-29 10-Q).

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
dgunning added a commit that referenced this pull request Oct 1, 2026
…s keep source offsets (5qqy, 6th2) (#1397)

5qqy: Filing._attachment_matches tested document.upper() in doc_type, a
substring match, so "EX-10.1" also selected EX-10.10..EX-10.19. Replaced by
_select_attachments, which sees the whole list: an exact filename wins,
then an exact document type, and the old substring match applies only when
nothing matched exactly, so "EX-10" still selects every EX-10 exhibit.
PR #1380 added only an exact-filename check ahead of the substring test,
which leaves the type collision in place. SeeQC's S-1
(0001213900-26-073222): "agreement" in EX-10.1 is 133 matches; the old
filter returned 378, adding EX-10.10/11/12.

6th2: the literal branch of _grep_text found positions in text.lower() and
sliced the original text with them. "İ".lower() is two characters, so each
one before a match shifted it by one. Literal search now runs
re.escape(pattern) with IGNORECASE over the original text. It keeps the
old scan's overlapping matches ("aa" in "aaa" is two) by searching again
from m.start() + 1 rather than using finditer; regex mode is unchanged.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
dgunning added a commit that referenced this pull request Oct 1, 2026
…n BDCs (1c6z) (#1398)

_bdc_portfolio fell back to find_bdc(identifier, top_n=1)[0] when the
identifier was neither a ticker nor a CIK. find_bdc scores every
whole-word hit near 100 (it only nudges toward the closest full name), so
the top hit identified nothing: "Golub" matches four Golub BDCs at
99.4-100 and silently became GOLUB CAPITAL BDC.

_pick_bdc_search_winner resolves a name only when (a) exactly one result
equals the query ignoring case, punctuation and a trailing corporate
suffix, or (b) the top hit scores >= 95 and the runner-up < 90. Otherwise
the tool returns AMBIGUOUS_BDC with up to five "name: identifier='CIK'"
suggestions (CIK, not ticker: non-traded BDCs have none).

The rule is PR #1380's (d12c057) with one change: it compared full names
literally, which would have made "Ares Capital" and "Prospect Capital"
ambiguous (siblings score 99.3; neither equals "...CORP"), though both
resolve correctly today. Ignoring the suffix keeps them resolving.
Measured on 17 queries: Golub, Ares, Blackstone, Blue Owl, Gladstone,
Oaktree, Monroe, Barings, Hercules -> ambiguous; Ares Capital, Main
Street, Prospect Capital, Hercules Capital, WhiteHorse, Saratoga -> one.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants