We do not want to review an AI writing tool by reading its feature page. We want to see what happens when we give it real content work.
For this Frase test, we created two different articles and followed the process from the first prompt to the final optimization results.
We saved screenshots during research, article planning, generation, SEO checks, GEO analysis and Improve. This lets us show the real workflow, including the parts that worked well and the parts that did not.
Try FraseAffiliate disclosure: this is an affiliate link. I may earn a commission if you sign up through it, at no extra cost to you.
Why We Created a Separate Frase Testing Methodology
A software review is much more useful when you know how the product was tested.
AI tools can look excellent in a short demo. A polished feature page can also make every function sound easy.
Real content work is different.
We wanted to know:
- how Frase handles a detailed prompt,
- how much research it does before writing,
- whether it changes the requested structure,
- whether the final article is complete,
- what happens to SEO and GEO scores,
- how useful Improve really is,
- whether a high score matches real content quality.
Our Frase Test Setup
We used two different content tasks.
AI Tools for Etsy Sellers
This test used a detailed commercial-style article prompt. We wanted comparisons, useful recommendations and information that could help an Etsy seller choose tools.
AI Content Creation Mistakes
This was a more natural educational blog post: “7 Things I Wish I Knew Before I Started Using AI for Content Creation.”
The two tests were intentionally different.
The first one tested a more structured, decision-focused article. The second tested a natural blog format that needed personality and clear teaching.
What We Measured
We did not judge Frase only by word count.
For each test, we looked at several parts of the process.
| Area | What we checked |
|---|---|
| Prompt understanding | Did Frase understand the topic, audience and requested structure? |
| Research | Did it analyze keywords, competitors, SERPs, questions and evidence? |
| Outline | Was the planned structure useful before generation? |
| Generation | Did Frase create a complete article? |
| SEO | Topic Coverage, Content Score, structure, keywords and metadata. |
| GEO | AI-search optimization signals and GEO scoring. |
| EEAT | Did the draft show enough authority, experience and trust signals? |
| Improve | Did the AI editing pass actually make the article better? |
Test 1: Best AI Tools for Etsy Sellers
Real Content Test #1Our first test started inside a Frase project.
We then gave Frase a detailed prompt for an article about the best AI tools for Etsy sellers.
This was not a one-line prompt.
We wanted to see whether Frase could follow a real content brief.
Frase Researched the Topic First
After the prompt was submitted, Frase started researching the topic.
This is one part of Frase that we wanted to document carefully.
The platform does not simply jump from prompt to finished text. Research happens before writing.
Frase Gave Us a Choice Before Full Generation
After analyzing the prompt, Frase gave us two options.
I like this because comparison data can be one of the most important parts of a commercial article.
Checking the table first gives the user a chance to catch problems before the whole article is written.
Frase Also Read Search Intent
This is important for a “best tools” article.
Someone searching for the best tools usually wants help making a choice. A general explanation of AI would not satisfy that intent.
The First Draft Was Generated
The initial draft was complete enough to review, but we did not stop there.
We wanted to see what Frase would find wrong with its own content.
Test 1: We Used Improve on the Draft
We opened the Improve workflow and let Frase analyze the article.
Frase then created comments from different specialist-style roles.
This was useful because the feedback was not only: “make the article better.”
Frase separated different problems.
Structure Specialist Suggested Key Takeaways
This was a reasonable recommendation for a long article.
SEO Specialist Suggested a More Specific Fix
Frase Also Suggested Supporting Evidence
What Test 1 Taught Us
What worked
- Frase followed a detailed content task.
- Research happened before generation.
- Search intent was part of the workflow.
- Improve found specific weaknesses.
- Suggestions were actionable.
What still needed humans
- fact-checking,
- checking statistics,
- judging recommendations,
- deciding whether suggested sections were useful,
- final editing.
Test 2: A More Natural Blog Article
Real Content Test #2The second test was different.
We asked Frase to create:
“7 Things I Wish I Knew Before I Started Using AI for Content Creation.”
This topic needed more than SEO structure. It needed a natural teaching style and a more personal feel.
Frase Suggested Its Own Article Setup
Then Frase noticed something interesting.
It decided that the “7 Things…” structure looked more like a listicle than a normal blog post.
This was a small but useful sign that Frase was interpreting the task instead of only repeating our settings.
Test 2 Had a Much Deeper Research Stage
Before writing, Frase researched several parts of the topic.
We could also see the recommended keywords.
Frase Researched Traditional Search and AI Search
This test also gave us AI-platform research.
We also found a difference between normal SERP competitors and AI visibility.
Frase Built a Detailed Outline
The outline was not just a list of seven headings.
Frase added talking points, examples and supporting evidence.
It also planned sections around Brand Voice and editing the first AI draft.
The Content Brief Was Surprisingly Detailed
Frase also created a broader content brief.
We liked the planning depth.
But we also noticed that Frase introduced some ideas of its own. That means the outline still needs human review before generation.
Then We Watched the Article Generate
Before generation, the editor showed:
- 0 words,
- 0% Topic Coverage,
- 0 headings,
- 0 links,
- 0 images.
During generation, the numbers updated live.
Later, the draft reached 993 words and 60% Topic Coverage.
Test 2 Final Generation Result
When Frase finished, the result looked very strong on the surface.
| Metric | Final generation result |
|---|---|
| Sections | 7 / 7 |
| Words | 2,275 |
| Topic Coverage | 100% |
| Headers | 7 |
| Links | 7 |
| Images | 0 at this stage |
This is exactly why we did not judge the article only by completion.
100% Topic Coverage Was Not the Same as 100% Quality
The finished draft had 100% Topic Coverage.
But its overall Content Score was only 68.
This became one of the most important findings in our Frase test.
Our Detailed Scores Were Even More Interesting
We then opened the detailed optimization panel.
| Metric | Result |
|---|---|
| Overall | 69 on the detailed panel |
| EEAT | 38% |
| GEO | 84% |
| SEO | 84% |
This result changed our view of the article.
It was not a bad draft. In fact, SEO and GEO looked strong.
But EEAT at 38% showed a clear weakness.
The article still needed more experience, stronger trust signals and human editing.
Then We Tested Improve on the Second Article
We did not stop after generation.
We opened Improve to see whether Frase could improve its own article.
GEO was broken into smaller areas such as:
- Quotability,
- Structure,
- Definitions,
- Takeaways,
- Citable Data.
The Improve Test Produced an Unexpected Warning
This is probably the most surprising result from the whole test.
Before applying the edits, Frase predicted that GEO could go down.
| Metric | Before | Predicted |
|---|---|---|
| Words | 2,275 | +36 |
| EEAT | 48 | 48 |
| GEO | 84 | 78 |
| SEO | 84 | 84 |
This is why our testing method includes the preview.
We do not assume that a feature called Improve improves every metric.
Final AFTER Result
After applying the changes, the article reached:
- 2,311 words,
- +36 words,
- Content Score 70,
- a new Key Takeaways section.
Test 1 vs Test 2
| Test 1: Etsy Tools | Test 2: AI Content Mistakes | |
|---|---|---|
| Main goal | Decision-focused commercial article | Natural educational blog post |
| Prompt | Detailed brief with table, pros/cons and FAQ | Natural “7 Things I Wish I Knew” article |
| Research | Topic and search intent research | Keywords, SERP, AI visibility, competitors and evidence |
| Improve test | Specialist comments and suggested fixes | Full score-based BEFORE/AFTER test |
| Strongest finding | Frase found useful structural and evidence gaps | 100% Topic Coverage did not mean publish-ready quality |
What We Learned From Both Tests
1. Frase Is Stronger as a Workflow Than as a Simple Writer
The value is not only the generated paragraph.
Frase connects research, planning, generation and optimization.
2. Research Is One of Its Strongest Parts
Both tests showed that Frase tries to understand the topic before writing.
The second test went especially deep with SERPs, keywords, AI-platform findings and evidence.
3. AI Can Create a Complete Draft
The second article reached 2,275 words and completed all 7 planned sections.
4. Complete Does Not Mean Publishable
This is our most important conclusion.
100% Topic Coverage looked excellent. EEAT at 38% told a different story.
5. Improve Is Useful, But Not Automatic
Frase found useful problems in both tests.
But the second test also showed a predicted GEO decrease. Every important change still needs review.
What We Did Not Use as a Success Metric
We did not judge Frase by:
- how fast it produced words,
- how impressive the interface looked,
- how many AI features were available,
- whether every score was green.
Those things can be useful, but they do not answer the main question:
Did the tool help create better content?
Our Frase Testing Rules
- Use real topics. We choose topics that could become real articles.
- Save the original prompt. This lets us judge whether Frase followed the request.
- Capture the workflow. We save screenshots before, during and after generation.
- Keep the scores. We record Content Score, SEO, GEO and EEAT when available.
- Test the editing tools too. A writer should not be judged only on the first draft.
- Look for unexpected results. We do not hide score drops or weak output.
- Do a human review. The final question is whether the content is actually useful.
How Our Frase Reviews Use These Tests
The two tests on this page are the evidence behind our larger Frase content cluster.
Instead of repeating every screenshot in one huge review, we use the same test data to answer different questions in more detail.
Frase Improve Content Review →
Frase SEO Optimization Review →
Want to Run the Same Kind of Test?
Use a topic you already know well. Save the prompt, check the research, generate the draft and compare the final article with what you would publish yourself.
Try FraseAffiliate disclosure: the Frase link above is an affiliate link. I may earn a commission if you purchase through it, at no extra cost to you.
Final Results: What Did Our Frase Tests Show?
Frase performed well as a research and content workflow tool.
It could take a detailed prompt, research the topic, build a plan and produce a complete draft.
The second test showed this clearly:
- 2,275 words,
- 7 of 7 sections,
- 100% Topic Coverage,
- SEO 84%,
- GEO 84%.
But it also showed the main limitation:
EEAT was only 38%.
Then Improve moved the article to 2,311 words and Content Score 70, but even that did not remove the need for human editing.
How We Tested Frase AI FAQ
How many Frase content tests did we run?
This methodology page focuses on two main content tests: an AI tools article for Etsy sellers and a natural blog article about mistakes people make when using AI for content creation.
Did we use real prompts?
Yes. We saved screenshots of the prompts and the generation process so we could compare the request with the final result.
Did Frase write a complete article?
Yes. In our second test, Frase generated 2,275 words and completed all 7 planned sections.
What was the Topic Coverage result?
The second test reached 100% Topic Coverage.
What Content Score did it get?
The completed article showed Content Score 68 on one final view and 69 on the more detailed optimization panel.
What were the SEO, GEO and EEAT scores?
In the detailed test, SEO was 84%, GEO was 84% and EEAT was 38%.
Did we test Frase Improve?
Yes. We tested Improve on both content workflows. In the second test, the article grew from 2,275 to 2,311 words and reached Content Score 70.
Did every score improve?
No. Before applying one set of changes, Frase predicted that GEO could decrease from 84 to 78 while SEO stayed at 84.
Do we recommend publishing the Frase draft directly?
No. We recommend checking important facts and sources, adding first-hand experience and doing a final human edit.
Where can I read the full Frase review?
Read our complete Frase Review for the overall verdict.