What can you build with Gemini 3.8 Flash that people will pay for?
SUMMARY
The best things to build with Gemini 3.8 Flash are narrow B2B agents that finish expensive, repeatable work and leave behind something a human can check. Repository maintenance and vertical RFP or security-questionnaire automation are the two strongest starting points.
The model is most compelling when the job has bounded source material, clear tools and a verifiable output. Its weaker computer-use and open-ended agent scores are a useful warning against selling the “autonomous employee” story too early.
The low API price is real, but the headline token rate understates what hard agent work costs. Gemini 3.8 Flash can burn substantially more reasoning tokens and tool steps than earlier models, yet its cost per successful professional task can still be far below premium alternatives.
Demand is already proven in the surrounding categories. Cursor, Glean, Sierra, Legora, Harvey and UiPath together represent more than $4.8 billion in reported annualized revenue or ARR, even though those figures should not be treated as a clean market-size calculation.
Coding is attractive, but another general AI editor is the wrong wedge. A maintenance agent that owns one recurring class of engineering debt, such as tested dependency updates or framework migrations, has a clearer job, buyer and success metric.
RFP and security-questionnaire automation may be the best bootstrapped B2B business of the group. Existing products already command five-figure annual prices, the work is document-heavy, model cost is tiny relative to deal value, and a human can approve the final answer before it goes out.
Finance and legal work are attractive for a different reason: the labor is expensive and the output can be audited against source material. Finance looks especially strong on current agent benchmarks; legal can support valuable document workflows too, but a 10% all-pass benchmark result makes serious human review non-negotiable.
Computer use is useful when it is boxed into repetitive portal work with visible state, retries and approval before irreversible actions. A general browser employee roaming across arbitrary software is still too unreliable to make a clean product promise.
Generic “chat with your company data” has become a weak standalone product. Retrieval is getting easier and cheaper, so the stronger move is to hide it inside a workflow that ends in a completed RFP, credit memo, redline, pull request or another deliverable people already budget for.
The best initial customer is a specialist B2B team that repeats the same costly task every week, not a mass-market consumer and probably not a giant enterprise on day one. The durable advantage should accumulate in customer data, integrations, accepted outputs and evaluation systems; if swapping Gemini for another model destroys the product, there was never much of a business around it.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →What actually changed with Gemini 3.8 Flash?
Gemini 3.8 Flash currently gives builders a combination that was much harder to buy cheaply before: near-frontier performance on several professional agent tasks, a one-million-token context window, multimodal input, built-in tools and very low API prices.
Google's current model documentation describes Gemini 3.8 Flash as production-ready and designed for long-horizon software engineering, autonomous agents and complex enterprise workflows. The same model can take text, images, audio, video and PDFs, return up to 65,536 tokens, call functions, execute code, search files, ground against Google Search or Maps, and use a computer in preview. Standard introductory pricing is $0.75 per million input tokens and $3.75 per million output tokens, with Batch and Flex priced at half those rates.
The jump from Gemini 3.7 Flash is large enough to affect product choices. Google's release evaluations improved across coding, finance, legal work, computer use and scientific reasoning, while Artificial Analysis independently puts the high-reasoning configuration on its cost-versus-intelligence Pareto frontier and measures output speed at roughly 281 tokens per second. Google also warns that the model tends to spend more reasoning tokens and tool calls on difficult jobs, so the token rate only tells part of the cost story.
We should still resist turning every improvement into a product idea. Gemini 3.8 Flash looks strongest when we give it a clear set of documents, clear tools and a result we can check. Its weaker scores show up on broad computer-use and open-ended agent tasks. That difference should decide what we build.
Is Gemini 3.8 Flash really good enough for work people pay for?
Gemini 3.8 Flash is already good enough for paid professional workflows when the work is narrow enough to verify, while a fully autonomous employee still goes beyond what the current benchmarks support.
The useful evidence comes from evaluations that look more like actual work. Google's current comparison puts Gemini ahead of the listed premium models on finance and legal-agent tasks, while the latest independent long-horizon coding leaderboard places it in the leading group. The legal evaluation is particularly demanding because the final work only passes when every required rubric item passes.
Broader agency is where our confidence should drop. Current computer-use and general-agent evaluations show a clear gap between Gemini 3.8 Flash and the strongest model. That gap should decide what we are willing to sell, because a polished demo can hide exactly the kind of failure that becomes expensive in production.
For us, the practical line is fairly clear today. Contract comparison, repository maintenance, financial workpapers and bid responses all give us something concrete to check. Letting Gemini roam through arbitrary software for hours asks much more from the model than the current evidence can support.
| Benchmark | Gemini 3.8 Flash | Claude Opus 5 | What it tells us |
|---|---|---|---|
| DeepSWE v1.1 | ~74% | ~74% | Same top cluster on long-horizon coding |
| Vals Finance Agent v2 | 61.4% | 58.6% | Strong fit for tool-using financial research |
| Harvey Legal Agent | 10.0% | 6.7% | Strong relative result, low absolute reliability |
| OSWorld 2.0 | 59.0% | 75.4% | Computer use still needs guardrails |
| Terminal-bench 4.0 | 19.1% | 51.8% | Open-ended agency remains a weak point |
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Is Gemini 3.8 Flash actually cheap once an agent starts thinking?
Gemini 3.8 Flash is genuinely cheap today, although difficult agent jobs can use so many extra reasoning tokens and tool steps that the real saving is smaller than the headline token price suggests.
Artificial Analysis found that the new model uses about 30% more output tokens per task than Gemini 3.7 Flash on its evaluation suite, pushing average task cost roughly 40% higher despite identical introductory token prices. Google has warned developers about the same behavior: Gemini 3.8 Flash "works harder" and may spend more tokens and take more agent steps to reach the answer.
DeepSWE lets us see the trade-off on the same coding harness. Gemini 3.8 Flash averages $2.36 per attempt at roughly 74% success, versus $11.84 for Claude Opus 5 at roughly the same success rate and $6.46 for GPT-5.6 Sol at about 73%. That works out to roughly $3.19 per successful Gemini attempt, around $8.85 for Sol and $16 for Opus 5. Gemini also uses about 143,000 output tokens and 166 agent steps there, compared with 60,000 tokens and 61 steps for Sol, so the low dollar cost comes partly from being able to spend a lot of cheap tokens.
Those economics look excellent when we charge for work worth hundreds or thousands of dollars. A cheap unlimited consumer plan is riskier because a few heavy users can run long agents all day. We would price around completed work or usage volume and watch cost per successful job much more closely than the token rate.
What are companies already paying AI agents to do?
For a Gemini 3.8 Flash business, the clearest demand today is in AI that changes code, handles legal and customer work, searches private company knowledge and completes business processes.
The current revenue figures are large enough to remove much of the speculation. Bloomberg reported Cursor above a $2 billion annualized revenue run rate. Glean says it has reached $300 million in ARR. Sierra says it entered its third year above $150 million ARR. Legora says it crossed $100 million ARR in less than 18 months after general launch. The Times has just reported Harvey at roughly $350 million ARR. UiPath's latest quarterly results put ARR at $1.938 billion, up 12% year over year.
Those six numbers add up to more than $4.8 billion in annualized revenue or ARR. The accounting definitions differ, so $4.8 billion would be a bad market-size estimate. As a rough check on what buyers are already funding, though, it is hard to ignore: AI-assisted coding, legal work, enterprise knowledge, customer operations and automation are already pulling in multi-billion-dollar spending.
The freshest UiPath numbers make the automation side especially interesting. Revenue rose 13% year over year in its latest quarter, and management says customers increasingly want AI agents combined with deterministic automation and governance. Sierra has moved in the same direction from another angle: its current positioning emphasizes paying for outcomes, and its Horizon agents are built to orchestrate interactions over days or weeks.
This is what buyers actually pay for: the pull request getting merged, the security review getting completed, the customer issue getting resolved or the legal matter getting processed. The model underneath can change while the job stays exactly the same.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Should you build a coding product with Gemini 3.8 Flash?
A coding product is one of the strongest Gemini 3.8 Flash opportunities right now, but the opening is much better in repetitive repository maintenance than in building another general-purpose AI editor.
Cursor already owns a lot of mindshare in the broad coding-assistant category and now charges $20 a month for Pro, $60 for Pro+, $200 for Ultra, and $40 to $120 per user for team seats. Its annualized revenue run rate above $2 billion shows that developers and companies will pay heavily when AI saves real engineering time. Competing head-on means reproducing an editor, agent environment, model routing, integrations, team controls and distribution that Cursor has spent years building.
A narrower maintenance agent has a cleaner reason to exist. We could connect a repository, detect stale dependencies, framework migrations, deprecations, flaky tests or repetitive security fixes, then make the change, run the test suite, inspect failures and open a pull request with evidence. Engineering teams routinely postpone that kind of work because it eats developer hours without shipping a visible feature.
As seen above, DeepSWE puts Gemini 3.8 Flash in the same top coding cluster as several premium models while its average attempt costs only $2.36 in that harness. The low cost lets us run a second reviewer, retry failed migrations, compare two approaches or test across multiple environments without destroying the gross margin.
We would pick one class of maintenance work and own it. "Keep our Python dependencies current and send tested PRs" is easy for an engineering lead to understand and budget for. The team can count merged PRs, hours saved, vulnerable packages removed and migrations completed, which ties value to measurable work instead of the chat experience.
Is RFP automation the best Gemini 3.8 Flash business?
RFP and security-questionnaire automation is probably the best bootstrapped B2B opportunity for Gemini 3.8 Flash because customers already pay five-figure annual prices for the workflow and the underlying AI cost can stay tiny.
Loopio's current entry plan starts at $20,000 per year and includes generative AI for RFPs, RFIs, due-diligence questionnaires, sales proposals and security questionnaires. Vanta says its Questionnaire Automation answers at least 80% of security questions automatically on average, with AI-generated answers accepted up to 95% of the time. Vanta also supports spreadsheet, document and third-party portal workflows, including a browser extension.
Gemini 3.8 Flash fits the job unusually well. A company can load policies, SOC reports, previous questionnaires, architecture documents and approved answers into File Search. Google currently charges for initial embedding creation and normal model tokens, while File Search storage and query-time embeddings are free. The agent can return structured fields containing the proposed answer, supporting evidence, confidence and any missing information.
The economics leave plenty of room for verification. A deliberately heavy request with 500,000 input tokens and 10,000 output tokens costs about $0.41 at today's standard introductory rate before extra tools. We could run multiple passes, check citations against source documents and still spend very little compared with a product that sells for thousands of dollars per year.
The market is already competitive, so "RFP software" is still too broad. A strong first product could handle security questionnaires for 20-to-300-person SaaS vendors, technical tenders for one industrial niche, or due-diligence questionnaires for one type of financial team. The customer's approved answer library, portal integrations and past accepted responses would make the product better every time it is used.
There is also an obvious next step once drafting works: move approved answers into the actual customer portal, attach the right evidence and stop when a new question cannot be substantiated. That brings Gemini's document reasoning and computer use together while keeping a human at the final approval point.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Can Gemini 3.8 Flash replace hours of finance research?
Gemini 3.8 Flash can already replace a meaningful chunk of repetitive finance research through auditable workpapers and screening work.
The strongest current evidence is Vals Finance Agent v2. Google's published run puts Gemini 3.8 Flash at 61.4%, ahead of Opus 5 at 58.6% and GPT-5.6 Sol at 53.8%. The benchmark involves tasks that can require web research, SEC filings, historical market data, calculations and source reconciliation, so it maps much more closely to analyst work than a generic finance Q&A benchmark.
A useful product would start with a company, fund, deal or credit and return a structured research file. It could collect filings and earnings materials, pull historical numbers, calculate margins and multiples, reconcile conflicting values and link every important figure to the underlying evidence. A private-equity analyst could use that for first-pass screening; a lender for a credit memo; a corporate-development team for an acquisition brief; an equity-research team for routine post-filing updates.
A finance team cares about where a number came from, whether the calculation can be reproduced and which assumptions changed. We can check all three. Predicting which asset will outperform leaves us with a much murkier product because the customer only learns whether the answer was good after money is at risk.
Finance already has large platforms such as AlphaSense and newer AI products such as Rogo, so we would pick one repeated deliverable and go deep. Updating a lender's recurring credit-review package or building a first-pass acquisition screen is much easier to explain, integrate and price than trying to become another full research terminal.
Can you really sell legal work powered by Gemini 3.8 Flash?
You can sell Gemini 3.8 Flash for legal document work today, but every serious product still needs lawyers reviewing consequential output.
Customers are already spending heavily on legal AI. Legora says it passed $100 million ARR less than 18 months after general launch and now serves more than 1,000 customers. The Times has just reported Harvey at about $350 million ARR with more than 200,000 lawyers using the product. Lately, large firms have also been putting more money into customized legal AI: the Financial Times reports that firms including Kirkland & Ellis and Freshfields are building or co-developing their own systems, with Kirkland committing $500 million to its platform.
Gemini leads Google's current comparison on Harvey's Legal Agent Benchmark, with a 10.0% all-pass rate. The benchmark covers more than 1,200 long-horizon tasks across 24 practice areas and grades the final work against expert-written rubrics. A 10% all-pass rate is impressive relative to rivals and still nowhere near the reliability a lawyer would accept without review.
That points us toward contract-heavy jobs where the source material is closed and the output can be reviewed. We could compare hundreds of supplier agreements against a playbook, extract obligations from leases, prepare a due-diligence issue list with links to clauses, build the first redline of a standard NDA, or identify where a new contract departs from previously approved fallback language.
Big law firms are also moving toward bespoke tools. That tells us firms will pay to encode their own templates, playbooks and previous work instead of relying entirely on a universal legal assistant. A smaller builder should go narrower again: one contract type, one practice workflow or one in-house process with an explicit lawyer review step.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Is Gemini 3.8 Flash computer use reliable enough to sell?
Gemini 3.8 Flash computer use is sellable now for supervised, repetitive portal work, while a general-purpose autonomous browser employee remains too unreliable for us to promise.
Google currently recommends Gemini 3.8 Flash for Computer Use and describes it as having high-accuracy UI interaction and reliable tool calling, although the capability still sits in preview. On OSWorld 2.0, Gemini scores 59.0% against 75.4% for Claude Opus 5. That is enough of a gap to keep human approval and recovery logic in the product.
We would therefore choose jobs where the page state is visible, the steps repeat and mistakes can be caught before an irreversible action. Think of updating supplier portals, transferring approved answers into customer security portals, maintaining catalog information across back offices, uploading compliance evidence or copying structured data between systems that never built a decent API.
Businesses are still spending heavily on automation. UiPath's latest results show $1.938 billion of ARR, up 12% year over year, while the company keeps launching agentic workflows across finance, retail, manufacturing and financial services. Its recent product direction mixes AI agents with deterministic automation and governance, which is close to the architecture we would want here.
We would verify important state changes, keep screenshots or logs, retry safe steps and ask for approval before submission, payment or another irreversible action. Those controls can turn imperfect computer use into a useful business process. Skip them and every navigation mistake lands directly on the customer.
Is "chat with your company data" still worth building?
Generic "chat with your company data" is a weak Gemini 3.8 Flash business today, even though private enterprise knowledge is clearly worth a lot of money.
Glean says it has reached $300 million ARR only 15 months after hitting $100 million, and it nearly doubled its Fortune 500 customer count year over year. That tells us companies care deeply about making internal knowledge usable by AI. It also tells us how high the bar has become: established platforms already handle permissions, connectors, search quality, enterprise security and agent workflows around that context.
Google has simultaneously made the basic plumbing easier. Gemini File Search imports, chunks, indexes and retrieves private documents, with free storage and free query-time embeddings. Retrieved text is simply billed as model context. A one-million-token window also lets Gemini 3.8 Flash work with very large collections once the relevant material has been retrieved.
A new product that stops at "upload files and ask questions" has very little defensible value. We would hide retrieval inside a job people already pay to finish. An RFP agent can pull approved security language to answer a buyer; a finance agent can retrieve prior deal work to update a memo; a legal agent can find fallback clauses before preparing a redline.
Google's current File Search limitation pushes us toward a proper workflow anyway. File Search cannot run in the same request as Google Search grounding or URL Context, so a product that needs private and public evidence has to retrieve them separately and reconcile the result. How well we handle that handoff can become part of the product quality.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Who should you sell a Gemini 3.8 Flash product to?
The best first customers for a Gemini 3.8 Flash product are specialist B2B teams that repeat the same expensive job every week and can tell whether we completed it properly.
Consumer products face a rougher equation. Gemini itself is widely available, generic AI subscriptions are crowded, and unlimited agent usage can become expensive when power users trigger long reasoning runs. Very large enterprises have the opposite problem: willingness to pay is enormous, but procurement, security reviews, integrations and deployment work can overwhelm a small company.
The most attractive first buyer sits between those two extremes. A 30-person SaaS company completing dozens of security questionnaires, a 100-person engineering organization carrying constant dependency debt, a boutique lender producing the same credit package every quarter, or an in-house legal team reviewing the same contract families already has people assigned to the problem. Saving even a few hours every week can justify hundreds or thousands of dollars a month.
Current software prices make that plausible. Loopio starts at $20,000 per year, while legal AI companies have reached nine-figure ARR by selling into high-cost professional work. Buyers in these markets already know that finishing an expensive job can be worth far more than another general AI seat.
The protection comes from what builds up around the model: approved customer data, deep workflow integrations, a history of accepted outputs and evaluations that catch bad work. If another model beats Gemini 3.8 Flash later, we should be able to route some tasks there without changing why customers buy the product. If swapping the model wipes out the advantage, the business was too close to the API.
What should you actually build with Gemini 3.8 Flash?
Yes, there are businesses worth building with Gemini 3.8 Flash right now, and our first choice would be a narrow agent that completes a valuable professional workflow from source material to a checked deliverable.
If we pick purely from what Gemini 3.8 Flash is best at, repository maintenance comes first. Coding benchmarks put the model near the frontier at low cost, while the size of the coding-agent market already shows that engineering teams will spend heavily when software gives them real developer time back. We would own one recurring class of engineering debt.
RFP and security-questionnaire automation may be easier to turn into a focused B2B company. Buyers already accept five-figure software pricing, Gemini is very good at reading large private document sets, and model costs are tiny beside the value of getting a sales or security process completed faster. A version built for one underserved type of buyer looks especially attractive.
Finance research comes next because Gemini's current Finance Agent result is unusually strong and the work can be made auditable. Legal teams may pay even more for document-heavy work, although current benchmark reliability makes human review non-negotiable. Portal automation has huge long-term upside as well, but we would keep the first version supervised until computer-use reliability moves higher.
Once we rank the opportunities, the pattern is pretty clear. Gemini 3.8 Flash gives us cheap reasoning, huge context, multimodal input and tools. The domain data, integrations, verification and precise definition of “finished” still have to come from us.
| Product to build | Who pays | Why it works now | Our ranking |
|---|---|---|---|
| Repository maintenance agent | Engineering teams | Frontier-cluster coding performance, low task cost, measurable PR output | 1 |
| Vertical RFP / security-questionnaire agent | B2B sales and security teams | Existing five-figure budgets, document-heavy workflow, easy human approval | 2 |
| Finance research / workpaper agent | PE, lenders, corp dev, research teams | Strong finance-agent result and highly auditable outputs | 3 |
| Legal contract / due-diligence agent | Law firms and in-house legal | Very high labor value and proven AI spending, with mandatory review | 4 |
| Vertical portal operator | Operations teams | Huge automation market and useful computer use, with current reliability limits | 5 |
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →OUR METHODOLOGY
The question behind this analysis is simple to ask and surprisingly easy to answer badly: what can you build with Gemini 3.8 Flash that people will actually pay for? Instead of relying on benchmark headlines, product demos or general AI enthusiasm, we broke it into the parts that matter commercially: what the model can do reliably, what hard agent work really costs, where buyers already spend meaningful money, how easily the output can be checked, and whether a narrow new product still has room to win.
For model capability, we gave more weight to evaluations that resemble paid professional work than to broad intelligence scores. Coding, finance, legal-agent and computer-use results were useful because they expose very different levels of reliability. A strong relative score was not automatically treated as production readiness; the 10% all-pass result on Harvey’s Legal Agent Benchmark is a good example of a result that is impressive against competitors but still demands human review.
For economics, we looked beyond the advertised token price. Gemini 3.8 Flash can spend more reasoning tokens and take more tool steps on difficult jobs, so we compared headline API pricing with observed task cost and, where possible, cost per successful attempt. That distinction is important when deciding whether a product should be priced per seat, per completed job or around usage.
For willingness to pay, we used current software pricing, reported ARR or annualized revenue, and product adoption as evidence that a workflow already has a budget behind it. We did not add those figures together and call the result a market size. Their role was narrower: to show where customers are already paying substantial amounts for coding, legal work, enterprise knowledge, customer operations and automation.
We then assessed each product idea point by point. Strong model fit counted more when the work ended in something verifiable; low inference cost counted more when the completed job was worth hundreds or thousands of dollars; and an incumbent was not automatically a negative if its pricing or revenue proved that buyers already care. The final ranking is a synthesis of those recent signals rather than a mechanical score.
For source quality, we prioritized first-party model documentation, benchmark owners, current product pricing and company financial disclosures, then used established reporting for private-company revenue figures and broader industry developments. We excluded unsourced social posts, recycled aggregation pages and claims that did not add a checkable number, benchmark result, product capability or commercial fact.
Key model and benchmark sources include Google’s Gemini 3.8 Flash launch and benchmark comparison, Google DeepMind’s Gemini 3.8 Flash model page, Google AI for Developers on the current model, Gemini API pricing, Gemini File Search documentation, Artificial Analysis’s independent Gemini 3.8 Flash evaluation, and Harvey’s Legal Agent Benchmark methodology.
Key commercial sources include Cursor’s current pricing, Bloomberg on Cursor’s annualized revenue, Glean’s $300 million ARR announcement, Sierra’s year-two revenue update, Legora’s $100 million ARR announcement, The Times on Harvey, UiPath’s latest quarterly results, Loopio pricing, Vanta Questionnaire Automation, and the Financial Times on bespoke legal AI.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Related blog posts
- What can you build with Gemma 4 that people will pay for?
- What can you build with GPT-6 Astra that people will pay for?
- What can you build with Amazon Quick that people will pay for?
- What can you build with Claude’s new computer use that people will pay for?
Who wrote this?
STEAL WHAT WORKS TEAM
We study profitable internet businesses, take them apart, and write down what actually works: pricing, distribution, growth, packaging. We turn 300+ proven examples into a database so founders can stop testing random ideas and start from proof. Explore the database →