An AI source list is a record of where each fact, quote, and claim in your work came from — which model or tool produced it, under what terms, and whether you're allowed to use the output the way you intend. If you want to use AI safely and ethically, the list needs four columns: the source, the license or terms, the verification date, and a note on what the source does not cover. That last column is the one most people skip, and it's the one that causes the most trouble.
The decision you're actually facing isn't "which AI tools are ethical." It's narrower and more annoying: when a vendor's terms page is silent on whether you can train on their output, or when a model's documentation doesn't say where its training data came from, what do you write down? The answer is a rule, not a research project. If the terms don't state it, it isn't permitted — you record that as a gap, not as a yes.
This site maintains an internal database of 360 AI tools, each with a pricing and capability snapshot recorded at verification time, most recently on 2026-09-24. That's a useful shape for a source list, and it's also a warning: snapshots go stale. A tool's terms page today may not match the terms page you read six months ago.
What are the four columns, and why does the fourth matter most?
Column one is the source itself — the tool, model, or dataset, named specifically enough that someone else could find it. "An AI chatbot" is not a source. "Claude, chat interface, output generated March 2026" is.
Column two is the license or terms of service, with the version date. Terms change. Recording which version you read is the difference between a defensible record and a guess.
Column three is your verification date — the day you actually read the terms, not the day you first used the tool. The 360-tool database mentioned above carries a verification date for exactly this reason: a pricing and capability note without a date is close to useless.
Column four is the gap note. What does this source not tell you? Where is the terms page silent? This is where ethical use gets decided, because silence is the default failure mode. A vendor who doesn't mention training on customer data hasn't granted you permission to assume they don't.
The unstated-means-permitted rule, and why it's the whole game
Most AI terms pages are written to grant narrow rights and stay quiet about everything else. That's not an accident — it's how legal documents manage risk. So you need a default rule to apply when you hit a blank.
The rule: if the terms don't explicitly permit a use, treat it as not permitted. Not "probably fine." Not "everyone does it." Not permitted.
Here's what that looks like in practice. Say you're writing a product description and you want to feed in a competitor's published spec sheet as context. The tool's terms say you retain rights to your inputs. They say nothing about third-party copyrighted material you upload. Under the rule, you don't upload the spec sheet — you summarize the relevant specs in your own words first, then paste those. Same output quality, no unresolved question in your source list.
That example is deliberately boring. The interesting failures are the ones where the gap is invisible until someone asks. If your source list says "used Tool X, terms silent on training data retention," you at least know what question to ask the vendor. If your source list says nothing, you find out during a legal review.
What actually goes in a model card or eval report
Some vendors publish documentation that answers most of your column-four questions before you have to ask. Knowing what a complete one contains tells you when you're looking at a real disclosure versus a marketing page with a "Responsible AI" heading.
A model card typically covers: intended use and explicitly out-of-scope uses, the training data's provenance and collection method, evaluation results broken out by task and demographic subgroup, known limitations and failure modes, and the date the card was last updated. An eval report goes further on the numbers — which benchmarks, what scores, and under what conditions the tests ran.
The tell is specificity. "Trained on a diverse dataset" is not provenance. "Trained on X, filtered for Y, with Z excluded" is. If a card lists evaluation scores but never names the benchmark, that's a gap you record in column four, not a fact you cite.
This matters more than it sounds. A model card that documents its limitations is giving you the gap notes for free. A card that doesn't is telling you the vendor either hasn't measured them or won't say — and both of those are things you want written down before you build a workflow on top of the tool.
Where this approach breaks down
Honest limits, because a source list isn't a compliance program.
It doesn't scale to every sentence. If you're logging the provenance of each clause in a 3,000-word draft, you'll spend more time on the list than the draft. The practical version is per-source, not per-claim: you record that a section was drafted with a given tool under given terms, not that sentence seven came from paragraph three of a specific output.
It can't resolve a vendor's silence for you. If the terms genuinely don't address your use case, the rule tells you to treat it as unpermitted — but it can't tell you whether the vendor would say yes if you asked. Sometimes the answer is to email them and get it in writing. That takes days, and for a one-off task it may not be worth it.
And it goes stale fast. That's the argument for the verification date column. A list you built a year ago and never revisited is a list of assumptions, not sources.
The source list isn't there to prove you were careful. It's there so that when someone asks a hard question, you already know the answer or already know it's a gap.
3 things people get wrong when they build one
- Logging the tool, not the terms version. "Used Jasper" tells you nothing about what you were permitted to do. "Used Jasper, terms v. [date], output used commercially" does.
- Treating a published model card as a green light. A card documents the model. It doesn't grant you rights to use the output. Those are two separate columns.
- Assuming silence means permission. This is the failure the whole list exists to prevent, and it's the one people default to because the alternative is inconvenient.
Key Takeaways
- An AI source list needs four columns: source, license or terms version, verification date, and gap notes.
- If a vendor's terms don't explicitly permit a use, record it as not permitted — silence is a gap, not a yes.
- Verification dates matter because terms and pricing change; a snapshot without a date is an assumption.
- A complete model card names training data provenance, out-of-scope uses, and benchmark-specific eval results.
- Log provenance per source, not per sentence — per-claim tracking doesn't survive contact with a real deadline.
The version of this that actually survives a busy week
Keep the list short and specific. One row per tool, four columns, updated when you re-read the terms rather than when you use the tool. For most people that's a handful of rows, not a database.
The habit that makes it stick is the gap column. Every time you read a terms page and can't find the answer to a question you actually have, write the question down. Over a few months you'll have a list of the specific things vendors won't say — and that list is more useful than any general guidance on AI ethics, because it's about your work.
If you're drafting content at volume and the prompt-writing overhead is what's eating your time, a zero-prompt generator like AI-Mind handles that layer by taking a description instead of a prompt. The source list still applies to whatever it produces.
Sources
- AI Tool Database, internal verified snapshot of 360 AI tools with pricing and capability records, 2026. Most recent verification date 2026-09-24.
- A Field Guide to AI Documentation: Model Cards, Eval Reports, Agent Cards, and More. What vendor documentation should contain, and how to read it.
- Is the best AI email writing assistant safe to use with confidential work emails?. Applying terms-of-service questions to a specific, high-risk use case.
- what is ai generated. Baseline definitions for tracking what counts as AI-generated material.
Frequently Asked Questions
Do I need a source list if I only use AI for brainstorming?
Yes, but a lighter one. Brainstorming output rarely ends up verbatim in your work, so the risk is lower — but if a brainstormed idea becomes a claim you publish, you still need to know where the underlying fact came from. Log the tool and the date you used it. If the idea survives into a draft, verify the fact independently before it ships.
What if a vendor's terms are genuinely ambiguous?
Record the ambiguity in your gap column and treat the use as unpermitted until you get clarity. If the task matters enough, email the vendor and ask for a written answer. Verbal assurances and support-chat replies are worth saving as screenshots, but they carry less weight than the published terms, which the vendor can change at any time.
How do I handle AI output that mixes my own writing with generated text?
Log it at the section level rather than trying to attribute individual sentences. Note which tool drafted which section and under what terms, then keep your own edits in a separate version if you need an audit trail. Trying to tag every clause is the fastest way to abandon the whole practice within a month.