How to Conduct an AI Charity Assessment: A Step-by-Step Guide for Funders
The phrase "AI charity assessment" tends to produce one of two reactions from grant-makers. The first is enthusiasm: AI can finally automate the evidence gathering that consumes hours of every funding cycle. The second is scepticism: AI makes things up, and a hallucinated charity report is worse than no report at all.
Both reactions are reasonable. Both miss the important point, which is that neither enthusiasm nor scepticism tells you how to do it. This guide is about the how: the specific steps involved in conducting an AI charity assessment that is accurate, consistent, defensible, and genuinely useful for the humans who make the final decision.
The workflow has six steps. Each step has a distinct purpose, and each depends on the one before it. Running them out of order — or skipping steps because the AI output looks good — is where most AI-assisted assessment processes go wrong.
Step 1: Define the scope of the assessment
Before any AI is involved, you need to be clear about what you are assessing and why. This sounds obvious, but it is frequently skipped in practice, and the consequences show up at the end of the process when trustees cannot agree on what the scores mean.
Scope covers four questions. Which charities are in scope — your current shortlist, a broader sector scan, a specific geographic cohort? What type of funding decision are you assessing for — a new grant, a renewal, a multi-year commitment, a first exploratory conversation? What is the relevant time horizon — are you assessing the charity as it is now, or its trajectory over the past three years? And what sources of evidence will you accept — only public data, or also charity-submitted materials?
Answering these questions before you start is not bureaucratic overhead. It is what allows the AI to gather relevant evidence rather than everything that is available, and what allows trustees to interpret scores in context.
Step 2: Set your criteria and scoring rubric
The second step is where most of the intellectual work happens. You define the criteria you will use, what each criterion means in specific terms, what evidence counts for each one, and how you will score what you find.
A standard AI charity assessment covers five to eight criteria. The most common are financial health and reserves, governance and trustee independence, impact evidence quality, leadership stability, safeguarding, and strategic or thematic fit. The right set for your foundation depends on what you fund and how you manage risk.
For each criterion, the rubric should specify: a definition concrete enough that the AI can apply it without interpretation; the evidence sources it should draw on; a scoring scale of typically one to five; and behavioural anchors for each score level — descriptions of what "good" and "poor" actually look like in terms of observable evidence, not just adjectives.
This step is harder to skip than it looks. Vague criteria produce vague AI output. "Financial sustainability" as a criterion will generate a different assessment each time because it means different things to different assessors. "The charity holds at least three months of unrestricted reserves and has shown stable or growing income over the past three years" will generate consistent output every time, because the AI knows exactly what to look for.
Read more about what a complete AI charity assessment template looks like if you are building your rubric from scratch.
Step 3: Run AI evidence gathering
With scope defined and the rubric in place, the AI evidence-gathering step can begin. This is where AI earns its value: systematically pulling evidence from the Charity Commission register, filed annual accounts, trustees' reports, Companies House records, and publicly available web sources for each charity on your list.
The key discipline here is source provenance. Every piece of evidence the AI surfaces should be traceable to a specific, checkable source. "The charity's most recent trustees' report, filed 14 months ago, states that reserves stood at £128,000 against annual expenditure of £340,000" is useful evidence. "The charity appears to have sufficient reserves" is not, because it cannot be checked. A well-structured AI charity assessment tool will attribute every finding to its source document so that the human reviewer can verify it.
The AI should also flag what it could not find. A charity with no filed accounts in the past 18 months is a different risk profile from one with complete filings. Absence of evidence is itself evidence, and the assessment record should reflect it explicitly rather than leaving blank fields.
Step 4: Apply AI scoring against the rubric
Once evidence has been gathered, the AI applies your scoring rubric to produce a structured assessment for each charity. It maps the evidence to each criterion, applies the scoring anchors you defined in step two, and produces a score for each criterion together with the evidence that drove it.
The output at this stage should be a structured document for each charity: criterion by criterion, with a score, the evidence used, and a brief explanation of how the evidence was mapped to the score level. This is what makes the assessment audit-ready. A trustee reviewing the output should be able to follow the reasoning from evidence to score for each criterion, and should be able to verify each piece of evidence independently if they want to.
Red flags — governance concerns, regulatory actions, very low reserves, missing safeguarding policies — should be surfaced separately and prominently, not buried in a criterion score. Some findings should be visible regardless of the overall score, because they require a specific response rather than just influencing a number.
Step 5: Apply human review
This is not optional, and it is not a formality. AI evidence gathering and scoring is fast, consistent, and comprehensive, but it does not have access to everything that matters. The assessor reviewing the output may know that the charity recently changed its CEO, or that a significant funder just withdrew, or that the impact methodology described in the report has been independently criticised. None of that appears in the public record. Human review is where contextual knowledge enters the process.
Effective human review is structured rather than open-ended. The reviewer should work through the AI assessment criterion by criterion, confirming scores where the evidence is clear, adjusting where they have additional context the AI did not have access to, and documenting the reason for any adjustment. This creates an audit trail that distinguishes "AI assessed" from "human adjusted and why," which matters both for governance and for learning over time.
Human review is also where qualitative judgement belongs: the phone call with the charity's finance director, the trustee's direct experience of the organisation's work, the professional view of a sector expert. These inputs cannot be automated, and they should not be. What the AI process does is ensure they are applied to a consistent, evidence-grounded foundation rather than to the memory of whoever happened to interact with the charity most recently.
Step 6: Document the decision
The final step is documentation — and it is worth treating as seriously as the assessment itself. A grant decision record that captures the criteria used, the evidence gathered, the scores assigned, and the final funding decision is one of the most valuable assets a grant-maker can build over time.
Decision records protect trustees against claims of inconsistency or bias. They allow new trustees to understand how previous decisions were made. They enable year-on-year comparison of assessments for returning grantees. And they create a feedback loop: when you can compare assessment scores against programme outcomes two years later, you start to learn which criteria actually predicted success and which ones did not.
An AI charity assessment that ends without a decision record has missed most of its long-term value. The time investment in documenting each decision is small relative to the value of the institutional memory it creates.
Common mistakes and how to avoid them
The most common mistake is running the AI without a rubric and using the output as if it were a structured assessment. It is not. Unstructured AI output is a research summary, not a scored assessment. It may be useful, but it cannot be compared across charities on a consistent basis.
The second most common mistake is treating AI scores as final without human review. AI evidence gathering is only as good as the public record, and the public record does not capture everything. Scores without human review are unfinished work.
The third is failing to document disagreements between the AI assessment and the human reviewer's judgement. When a trustee overrides an AI score, the reason should be written down. That is where institutional learning comes from.
What this looks like in practice
ClearGiving implements all six steps in a single platform: a customisable rubric that the foundation defines once, automated evidence gathering from the Charity Commission and public sources, AI scoring against your criteria, structured human review tools, and a decision record for every assessment. You can read more about the underlying methodology, or review the charity due diligence checklist for the specific evidence sources used at each step. If you want to see the process in action on your own shortlist, request access to the platform.