Files
CIAgent/apps/api/app/prompts/report_generation.py
T
sakshamandClaude Sonnet 5 18305b545c Backfill sparse report sections, add enrichment section refresh, and bootstrap first-admin
Reports: the LLM reliably used company_enrichment for prose fields but
inconsistently populated the parallel Finding-list/string-list fields from
the same evidence, even with progressively more explicit prompting. Add a
code-level backfill (products, recent developments, financial signals,
strategic initiatives, regulatory signals, risks/opportunities mirrored
from SWOT, unknowns, monitoring recommendations) that only ever fills in
what the model left empty, never overwrites what it produced.

Enrichment tab: reorder sections (Products/Recent updates before
Customers/Competitors) and add a per-section "Refresh" button that
re-fetches just one of NinjaPear's six independent per-company endpoints
when it came back empty - confirmed live that a data-coverage gap (e.g.
Amazon returning no products) is real provider behavior, not a bug.

Auth: the first account registered on a deployment with zero existing
admins is now auto-promoted to admin, closing the chicken-and-egg gap
where the only path to admin access was direct DB access. Self-heals if
the last admin ever deletes their account.

Also bumps nginx's proxy_read_timeout for api.ciagent.org to cover the
enrichment refresh's synchronous funding-endpoint call (up to 5 minutes
per NinjaPear's docs).

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-08-06 17:41:02 -04:00

185 lines
10 KiB
Python

"""Task D: report generation. Turns accumulated evidence (source documents +
detected changes) for a company into the structured CI report (spec section
17). The LLM only ever sees evidence this pipeline collected - it is
explicitly instructed not to introduce outside knowledge, and every
non-trivial claim must carry a confidence label and evidence references.
"""
from __future__ import annotations
from pydantic import BaseModel, Field
from app.analysis.llm.base import LLMProvider
from app.prompts.base import build_user_prompt
from app.prompts.schemas import ConfidenceLabel, EvidenceRef, Finding
SYSTEM_PROMPT = (
"You are a competitive intelligence analyst producing a structured report about a "
"company, using ONLY the evidence provided - the company_profile block (discovered once "
"at onboarding from the company's real website and search results, not your own "
"knowledge), the company_enrichment block (when present - fetched once from a paid "
"third-party data provider at onboarding, also real evidence, not your own knowledge), "
"stored source documents, and previously detected changes. Never introduce "
"facts from outside knowledge, even if you recognize the company - if it isn't in the "
"evidence block, it doesn't go in the report.\n\n"
"company_profile and company_enrichment are real evidence and must ground every "
"applicable section even when source_documents and detected_changes are sparse or "
"empty - do not say 'insufficient evidence' for a field company_profile or "
"company_enrichment already answers. In particular, the list-of-Finding sections are "
"NOT limited to newly detected changes - they must also represent the company's "
"current, confirmed state whenever company_enrichment/company_profile answers them "
"directly, the same way you would use that evidence for a prose field:\n"
"- products_and_services: if company_enrichment.products is present, emit ONE Finding "
"per product (headline = product name, summary = its description/category, "
"confidence = confirmed). A Finding's evidence array may be empty ([]) - "
"company_enrichment is itself the evidence, and an empty evidence array is valid and "
"expected here since there is no source_document or url to cite for it. For example, "
"given company_enrichment.products containing {\"name\": \"Acme Pay\", \"category\": "
"\"Payments\", \"description\": \"Lets merchants accept cards online.\"}, emit: "
"{\"headline\": \"Acme Pay\", \"summary\": \"Lets merchants accept cards online "
"(Payments).\", \"confidence\": \"confirmed\", \"evidence\": []}. Do not leave "
"products_and_services empty just because no source_document specifically announces a "
"product, and do not skip a product merely because you have nothing to put in its "
"evidence array.\n"
"- recent_developments: if company_enrichment.recent_updates is present, emit ONE "
"Finding per genuinely distinct update (headline paraphrasing its text, date = its "
"date, confidence = confirmed since it is the company's own published content, "
"evidence = [{\"url\": its url, \"description\": \"one sentence\"}]). You may skip "
"near-duplicate or routine items, but do not leave this empty when recent_updates has "
"substantive entries.\n"
"- strategic_initiatives and key_inferred_projects: these are your own synthesis "
"across company_enrichment.recent_updates, products, and description - identify "
"recurring or notable strategic themes (a new market entered, a technology bet, a "
"business-model shift) that no single item states outright. Use a lower confidence "
"label here (likely/possible) since this is inference, not a directly-stated fact - "
"an empty evidence array is fine here too. Do not leave these empty just because "
"nothing states 'this is a strategic initiative' in so many words - if you can name a "
"theme in executive_summary or market_positioning, that same theme belongs here as a "
"Finding/InferredProject too, not only as prose. If market_positioning or "
"company_overview names ANY theme, direction, or bet the company is making, restate it "
"here as at least one Finding/InferredProject before considering this section done.\n"
"- regulatory_and_legal_signals: derive from anything in company_enrichment or "
"source_documents touching licensing, compliance, jurisdictions of operation, or legal "
"structure (e.g. a payments company operating in many countries implies licensing/"
"compliance obligations even if no single document states them). If genuinely nothing "
"in the evidence touches this even indirectly, emit one Finding with confidence "
"insufficient_evidence explaining that rather than an empty list.\n"
"- unknowns_and_missing_data and monitoring_recommendations: these two are NOT findings "
"about the company - they are your own meta-analysis of this report and evidence set, so "
"they almost never have 'insufficient evidence' as a valid reason to be empty. You "
"always have something to say: unknowns_and_missing_data should name specific "
"categories of information this evidence set does NOT cover (e.g. 'no financial "
"statements or revenue figures were available', 'no employee reviews or Glassdoor "
"sentiment data', 'pricing details were not provided'); monitoring_recommendations "
"should name specific things worth watching for on the next run (e.g. 'watch for "
"updates to the products list', 'monitor for new funding rounds', 'track leadership "
"page for executive changes'). Leave these empty ONLY if you truly cannot think of a "
"single gap or follow-up, which should be rare.\n\n"
"Every finding must still be traceable to the evidence given and must carry an honest "
"confidence label: confirmed (the source states it directly), strongly_indicated, "
"likely, possible, unconfirmed, or insufficient_evidence. When evidence is genuinely "
"thin or missing for a section, say so explicitly in that section rather than "
"inventing content. Distinguish clearly between what a source states and what you are "
"inferring."
)
class InferredProject(BaseModel):
project_name: str
status: ConfidenceLabel
summary: str
confidence: float = Field(ge=0.0, le=1.0)
evidence: list[EvidenceRef] = Field(default_factory=list)
signal_types: list[str] = Field(default_factory=list)
alternative_explanations: list[str] = Field(default_factory=list)
class SwotAnalysis(BaseModel):
strengths: list[str] = Field(default_factory=list)
weaknesses: list[str] = Field(default_factory=list)
opportunities: list[str] = Field(default_factory=list)
threats: list[str] = Field(default_factory=list)
class ReportContent(BaseModel):
executive_summary: str
company_overview: str
products_and_services: list[Finding] = Field(default_factory=list)
market_positioning: str
recent_developments: list[Finding] = Field(default_factory=list)
strategic_initiatives: list[Finding] = Field(default_factory=list)
key_inferred_projects: list[InferredProject] = Field(default_factory=list)
leadership_changes: list[Finding] = Field(default_factory=list)
hiring_signals: list[Finding] = Field(default_factory=list)
technology_signals: list[Finding] = Field(default_factory=list)
patent_signals: list[Finding] = Field(default_factory=list)
manufacturing_and_expansion_signals: list[Finding] = Field(default_factory=list)
partnerships_and_acquisitions: list[Finding] = Field(default_factory=list)
financial_signals: list[Finding] = Field(default_factory=list)
regulatory_and_legal_signals: list[Finding] = Field(default_factory=list)
customer_sentiment: str
competitor_comparison: str
swot: SwotAnalysis
risks: list[str] = Field(default_factory=list)
opportunities: list[str] = Field(default_factory=list)
unknowns_and_missing_data: list[str] = Field(default_factory=list)
monitoring_recommendations: list[str] = Field(default_factory=list)
methodology: str
limitations: str
async def generate_report(
llm: LLMProvider,
*,
company_name: str,
company_aliases: list[str],
competitors: list[str],
monitoring_focus: str | None,
industry: str | None,
documents: list[dict],
detected_changes: list[dict],
sources_failed: list[str],
description: str | None = None,
official_website: str | None = None,
headquarters: str | None = None,
country: str | None = None,
region: str | None = None,
public_identifiers: dict[str, str] | None = None,
enrichment: dict | None = None,
) -> ReportContent:
evidence = {
# Discovered once at onboarding (app/services/discovery_service.py)
# from the company's real website + search results - genuine
# evidence, not the LLM's own background knowledge, and the only
# evidence available before any monitoring run has collected
# source_documents/detected_changes.
"company_profile": {
"name": company_name,
"description": description,
"official_website": official_website,
"aliases": company_aliases,
"competitors": competitors,
"monitoring_focus": monitoring_focus,
"industry": industry,
"headquarters": headquarters,
"country": country,
"region": region,
"public_identifiers": public_identifiers or {},
},
# Fetched once, at onboarding, from a paid third-party provider
# (NinjaPear) - see app/services/enrichment_service.py. Absent for
# any company created before this feature existed, or whose
# enrichment never completed successfully - never fabricated.
"company_enrichment": enrichment or {},
"source_documents": documents,
"detected_changes": detected_changes,
"sources_that_failed_to_collect": sources_failed,
}
user_prompt = build_user_prompt(
"Produce a full structured competitive intelligence report from this evidence. "
"If a section has no supporting evidence, say so explicitly instead of leaving it "
"generic or inventing content.",
evidence,
)
return await llm.generate_structured(SYSTEM_PROMPT, user_prompt, ReportContent)