Blog
Survey Text Analysis

How to Choose Verbatim Coding Software

Summary
  • Get clear on your own requirements before comparing vendors
  • Know what to look for: quality, speed, ease of use, control and flexibility, privacy, security, and cost
  • Test the tool on your own data before you buy

If you run open-ended questions in your surveys, you already know the tradeoff. Open-ends give you the "why" behind the numbers, the honest, unprompted feedback that a rating scale just can't capture.

The catch is that turning hundreds or thousands of those free-text responses into something you can actually report on is a real project, not something you knock out in an afternoon.

That's what verbatim coding software is for. AI is genuinely reshaping how this work gets done, and plenty of tools now claim to do it well. So how do you know which one to choose?

As the founder of a company that develops AI software for analysing open-ended survey responses, I’ve spoken with hundreds of researchers and buyers trying to make this exact decision.

I’m obviously biased, so keep that in mind as you read. But those conversations, along with what we’ve learned from our customers, have given me a pretty good view of what researchers actually care about when evaluating these tools, what you should look for, and what questions are worth asking before committing to a platform.

What Is Verbatim Coding Software?

Verbatim coding software helps you turn open-ended survey responses, sometimes called verbatims, into structured themes (or "codes") that you can count, compare across segments, and report on. It's a two-step process: the software builds a codebook (the set of themes), then applies it across your responses, so instead of reading and tagging every single response by hand, you end up with a reliable coded dataset much faster.

For a fuller look at how the process works and why teams move away from doing it manually, our guide to verbatim coding covers that in more depth.

What Should You Consider About Your Own Needs First?

Before comparing vendors, make sure you know what you really need. A tool that works well for one team might not suit another, often just because their data and research process looks different.

A few questions worth answering first:

  • What's the volume of your coding?
    If you're only dealing with a handful of responses, tens or low hundreds, you can often manage manually or with a simple tool. Once you're regularly coding larger volumes, you need something built for scale, or the time savings you're hoping for won't materialize.
  • Is this a one-time project, or something you do regularly?
    A single one-off project, a student working on a thesis, for example, might be fine with a general-purpose tool like ChatGPT or Claude. If you're a professional researcher coding data as an ongoing part of your work, you need a dedicated, purpose-built tool that holds up project after project.
  • How important is quality to you?
    Are you willing to compromise on quality for a lower price, or is it worth paying more to get it right? Knowing this upfront means you're evaluating every tool against the right bar, rather than being swayed by whichever one happens to be cheapest or flashiest.
  • Who's actually going to use it?
    How tech-savvy is your team, and how much training or onboarding will new people need? Just as important: can you trust the tool to produce stable, consistent results across different team members, not just whoever originally set it up?
  • What format is your raw data in, and what happens after coding?
    Coding usually isn't the last step, so check whether the tool can import and export data in the formats your workflow actually needs, and whether it fits smoothly into the rest of your stack.
  • Do you already have a codeframe from a previous wave of this research?
    If so, you'll want a platform that can work with what you've already built, so you can stay consistent with past waves instead of starting from scratch.
  • Do you have multiple languages in your data?
    If you run a global tracker, you need a tool that can code consistently across languages, preferably without requiring a whole global team to support it.
  • Is privacy and security a concern for you?
    Are you working under GDPR, HIPAA, or similar requirements? This alone can rule tools in or out before you even get to comparing features.
  • What's your budget?
    Are you working with a budget that supports a professional, dedicated tool, or is this a smaller, cost-constrained project? Can you commit to an annual subscription or you need a pay per project option? 

Answer these first, and the rest of your evaluation goes a lot faster, because you'll already know which features actually matter for your situation and which ones are just nice extras.

What Questions Should You Ask a Vendor Before You Commit?

It’s always great to start with a demo, but it’s also important to ask the right questions. The questions below can help you understand whether a text analysis tool is a good fit for your team, your data, and the way you actually work.

  1. How accurate is the coding compared with human coding? How does the platform handle new topics, or different ways people say the same thing, slang, emojis, typos, and so on? 
    Traditional text analysis platforms were limited in what they could understand, but newer AI-based tools can now reach a level of quality comparable to human coding. You want a platform that can truly understand nuance and context, and adjust to new topics and your specific niche without you having to spend time teaching it first.
  2. How does it fit into your existing research workflow?
    Make sure you can easily import and export data in the formats you already use, so the platform fits into your existing process rather than creating extra work.

    If you run tracking studies, check whether the tool can work with an existing codeframe. That way, you’re not starting from scratch or losing the ability to compare results over time between waves.
  3. How much time and effort does it actually take to get to a final, reportable coded dataset?
    Check how long it takes to generate a good codebook, how well it covers your data, and how many iterations are typically needed before you have a final coded dataset.

    AI can produce results that look plausible at first glance, but the real test is how much review, editing, and rework is needed before those results are ready to use.
  4. How is pricing structured, how predictable is it, and can it fit your business model?

    You want pricing that’s easy to understand and budget for in advance, as well as a model that actually works for how you operate. Many research agencies work project by project, so flexibility can be important. For brands or teams with consistently high volumes, volume-based pricing or discounts may matter more.
  5. Can you test it on your own data?
    A demo is useful, but there’s no better way to evaluate a platform than trying it with your own data and seeing how it fits your actual process and workflow.

    A platform that’s easy to use and quick to set up should be happy to offer trial access.

What Should You Look For in Verbatim Coding Software?

Three things matter most here: coding quality, speed, and ease of use. Beyond that, there are several other factors worth considering, including how much control you have over the final output, integrations with your existing workflow, multilingual support for global research, privacy and security, and the level of support available.

Accuracy and coding quality

Accuracy is where vendor claims and reality tend to drift apart, so it's worth checking a few things directly rather than taking a pitch at face value. Code a small sample of your own data by hand and compare it to the tool's output, you're not looking for a perfect match, just checking that the differences make sense and are easy to fix if needed.

Nuance is where the real gap shows up. More traditional platforms are often keyword-based, matching on specific words and phrases, which means they tend to struggle with sarcasm, typos, and slang. This used to be a limitation across the board, older, machine-learning-based platforms could never quite match human-level judgment on context and tone. 

Large language models changed that, since they work at the level of meaning rather than exact wording. Take a response like "oh great, another broken feature", a keyword-based tool might spot the word "great" and code it as positive, while a meaning-based tool recognizes the sarcasm and codes it correctly as negative, much closer to how a person would actually read it.

Then there's coverage: how much of your data is left over once the tool's done its pass. On many platforms, a meaningful chunk of responses, sometimes 20-30%, ends up dumped into a catch-all "other" bucket that still needs coding by hand. 

A well-built codebook should realistically cover 90-95% of your data on its own, with most agencies aiming to keep unclassified "other" responses down around 5-10%. 

This is genuinely hard to get right (it was true for Blix too, in the early days, before real investment went into making the codebook fit each dataset well out of the box), so ask a vendor directly what percentage typically still needs manual review, rather than just what the tool claims to automate.

Speed

Ask how long it actually takes to go from raw responses to usable, coded data you can actually report on, not just how long the AI takes to run, but the full process including review. 

The review step is on you, not the tool, so if the initial coding quality is low, that time comes straight out of your schedule. Ask the vendor how long it should take for a dataset your size, and better yet, run an actual trial to see how long it really takes on your own data.

Ease of use

Building software tools that are intuitive and easy to use is not an easy task.If a platform takes weeks to learn, or needs one dedicated person on your team just to run it, you're paying for that extra time on top of the tool itself. 

Look for something your whole team can actually pick up and use, not just the one person who set it up. A simple test during a demo is to ask yourself at the end: could I use this on my own right now?

At Blix, the most common reaction we hear after a demo is some version of "wow, this looks so easy". That’s the kind of feeling you should be looking for..

Results that you can verify

If a stakeholder asks what people actually said, you should be able to answer right away instead of digging through a spreadsheet. Being able to click into any theme and see the exact responses behind it is crucial for trust and traceability.

This matters even more with AI involved: you want zero chance of hallucination, every code and every theme should be 100% grounded in your actual data, with a direct line back to the source. AI can suggest the coding, but a person should always be able to trace it back and verify it.

Control and flexibility

The platform is there to do the heavy lifting, but you should always be able to stay in the driver's seat and step in when you need to, since only you have the full business context behind the data. 

Some platforms hand you a finished codebook and that's it, no room to adjust. It might be a perfectly good codebook that just isn't the best fit for how you want to run your research. 

You should be able to split, merge, or delete codes, and Blix specifically also lets you add plain-language instructions to fine-tune how the tool interprets a specific topic.

If you disagree with how something's been categorized, or your client or stakeholder asks for a change, you should be able to make that adjustment yourself, at any point in the flow.

There's also the question of how feature-rich the platform is versus how simple. Some tools pack in a wide range of research features; others deliberately do one thing and keep everything else out of the way. Neither is automatically better, it really comes down to whether you actually need that extra range, or whether it just adds friction to a task that should be quick.

Worth checking too: is this a platform that was genuinely built for AI, or one that had AI added on top? This isn't a minor detail, 67% of research professionals now say AI capability is a critical or key factor when choosing a research vendor, so it's worth knowing which kind of platform you're actually looking at.

In practice, this comes down to how much of the coding still lands on you. A platform that had AI bolted onto an older, manual-first system tends to only automate part of the job, so you're still doing a lot of the coding work yourself, just with some assistance. 

A platform built around AI from the start is designed to do most of the coding itself, with your role shifting to reviewing and refining its output rather than producing it from scratch. Same end goal, very different amount of your own time involved.

Finally, ask whether the tool needs retraining and calibration for every new project or topic. Some platforms have to be set up from scratch for each new survey subject or industry, which only really pays off on large, ongoing studies where that setup investment makes sense. 

A tool built on a large language model shouldn't need that: it already has broad world knowledge, so it can handle a brand-new niche, agriculture, healthcare, transportation, whatever your survey happens to cover, without needing to be trained on it first.

Reporting and integrations

Theme frequencies are a good start, but most teams need more: segment cuts, tables they can drop straight into a client deliverable, outputs that actually match how you report.

Coding is usually not the final step in your workflow, so check if you can easily export coded data to the next tool, like Excel, SPSS, or a reporting deck. 

The real issue usually isn't whether you can export a file, it's whether it actually fits the next tool in your process. If you end up spending time rearranging columns or copy-pasting pieces around just to get the data into a usable shape, that's a real cost, even if the coding itself was fast.

And if you're running tracking studies with recurring waves, ask how the tool handles reusing and updating a codeframe over time, and whether definitions stay consistent as new data rolls in.

Reporting and integrations

Theme frequencies are a good start, but most teams need more: segment cuts, tables they can drop straight into a client deliverable, outputs that actually match how you report.

Coding is usually not the final step in your workflow, so check if you can easily export coded data to the next tool, like Excel, SPSS, or a reporting deck. 

The real issue usually isn't whether you can export a file, it's whether it actually fits the next tool in your process. If you end up spending time rearranging columns or copy-pasting pieces around just to get the data into a usable shape, that's a real cost, even if the coding itself was fast.

And if you're running tracking studies with recurring waves, ask how the tool handles reusing and updating a codeframe over time, and whether definitions stay consistent as new data rolls in.

API and integrations

If you might need it, ask about API access. It's a bit more advanced, but very much worth it, especially with high volumes of textual feedback. At Blix we offer different types of text analysis APIs:

Batch coding 

This is the common one. Say you want to analyze your data once a day, once a week, or once a month, you just send your responses through the API and get them back coded automatically. It can plug straight into a dashboard or your survey platform, so the coded results just show up wherever you already work, no one has to run anything manually.

If you're running the same tracker every month or quarter, this can also mean the whole thing just runs itself, no manual work at all.

Live coding

This is still a relatively new capability, and not many platforms offer it yet. Traditionally, coding was an offline process: you waited for responses to come in, exported the data, and then coded everything in batches.

With real-time coding, each response can be analyzed moments after the respondent clicks “Next” on the survey question.

That opens up completely new use cases. You can trigger follow-up questions based on what someone just wrote, create surveys that adapt to free-text responses, update dashboards as data comes in, or immediately flag urgent feedback, such as an angry customer complaint or a health and safety issue, so the right person can act on it right away.

Multilingual support

If you run global research, check whether the tool can code across multiple languages without a separate translation step. 

At Blix, this is one of our differentiators: you can have one codebook, built in English (or any other language), and use it to code across a mix of 40 different languages out of the box, without translation.

Data privacy and security

Data privacy deserves careful attention. Survey responses can include sensitive information, so  it's worth understanding where your data is processed and stored, how long it’s kept, and if it’s ever used to train models outside your account. 

Uploading client feedback straight into a general-purpose AI chatbot can risk making your data public. A purpose-built research tool should run in a secure, private environment where nothing you upload trains anything outside your account.

Service and support

Last but not least, don’t overlook support. - Find out how you actually reach someone when you have a question, and how responsive and helpful the team tends to be. It's much better to know that before you're relying on a tool mid-project than to find out the hard way.

How Do You Test a Vendor on Your Own Data Before Buying?

One of the most useful things you can do before committing to a platform is test it with your own real data and experience the full workflow for yourself.

Ideally, choose a survey that has already been coded, run the same responses through the platform, and compare the results with the coding you already know and trust.

There are two useful steps to this comparison:

Step 1: Compare the codeframe 

Let the platform generate its own codebook from your data, then compare it to the one your team built manually. Keep in mind you'll be working with the platform along the way, not just receiving a finished product, so also pay attention to how easy it is to get from the platform's initial suggestion to a codebook you're actually happy with.

Step 2: Compare the coding itself

Take the exact codebook your team used for the manual coding and load that same codebook into the platform. Now you're comparing how the platform codes each response against how your team coded it, using the identical codebook, which isolates the actual coding accuracy from any differences in codebook design.

Statistical reliability measures, like Cohen's Kappa, are the standard way to score that comparison objectively instead of just eyeballing it.

It's worth remembering that even the same human coder, coding the same data twice, won't hit 100% agreement with themselves. Coding involves judgment, and there isn’t always one objectively correct answer, so you should expect some differences. 

The goal is not perfect agreement, but a consistently high level of alignment with the coding you already trust. . 

The real test is simpler than a statistic, though: would you feel comfortable putting these results in front of a client and standing behind them?

As a reference point, Blix’s internal testing shows 95%-99% agreement between Blix’s AI coding and human coders across the datasets we've tested (we'll be publishing the full methodology behind that in a separate piece).

What Should You Avoid When Choosing Verbatim Coding Software?

A few types of tools are worth going into with your eyes open, or steering clear of altogether.

Tools still built around manual coders, with AI added on top

A lot of platforms in this space were originally designed for human coding teams, and have layered some AI features on since to look current. 

In practice, that usually means long hours of "training" the tool and manual cleanup on your end, not because AI can't do better, but because the platform was never actually built around AI doing the heavy lifting.

Keyword-based or older machine-learning tools

These generally can't match the quality of newer, large-language-model-based tools, especially on anything involving nuance, sarcasm, or context. 

Overloaded, overly complicated platforms, or tools where coding is just a side feature

Some tools pack in far more features than you'll ever use, with a UI that takes real time to learn. Others are built primarily for something else entirely, a survey platform, a CX suite, a broader research tool, with open-ended coding tacked on as a minor add-on rather than the core product. 

Worth being honest about whether you actually need all that extra range, or whether it's just getting in the way of the one thing you need done well.

Tools that aren't private and secure

Ask where your data goes, how long it’s stored, and whether it’s ever used to train models outside your account. Make sure the terms of use explain this clearly.

How Does Verbatim Coding Software Pricing Typically Work?

Verbatim coding software pricing usually comes down to a few common structures:

Pricing model What it means
Per-response (per-verbatim) pricing Priced per coded response rather than per token or per code applied, the easiest model to budget against since you can calculate your maximum spend from your expected survey volume upfront. For example, if your survey generates 1,000 responses, you know upfront you're paying for 1,000 coded responses, no matter how long each one is or how many codes it ends up with
Token or credit-based pricing Cost varies by response length, language, or number of codes applied per response; harder to predict your total cost in advance
Pay-as-you-go Usage-based, billed monthly for what you actually consume. No commitment, flexible for project-based work or agencies with volatile month-to-month volume, and a good way to test a new tool before committing further
Annual subscription A bundled package of credits; larger packages bring the per-unit cost down significantly at higher volume

Per-response pricing is easy to understand with an example. If you have 500 survey respondents and 3 open-ended questions, that’s 1,500 responses to code. If a vendor charges per response, your maximum cost is 1,500 times the price per response, so you know the total before you field the survey. 

This is helpful for agencies making client proposals, since you can include the cost upfront. Pricing models based on response length, number of codes, or token usage are harder to plan for because you often don’t know the final cost until the project is finished.

It’s also worth thinking about whether a subscription or pay-as-you-go model better fits the way you work. 

If your project volume is unpredictable, pay-as-you-go pricing gives you more flexibility and lets you pay only for what you use. It can also be a good way to try a platform on real projects before making a larger commitment.

If you have a stable, high volume of text to analyze, a subscription or volume-based plan may be more cost-effective.

Also ask if the price changes for translation or re-running an analysis, since some platforms charge extra for these.

Why Blix

We built Blix to help researchers unlock the value hidden in open-ended feedback, without the time and effort manual coding usually takes. 

It delivers high-quality, human-like coding in minutes, through an intuitive easy to use platform.

Blix reads for meaning rather than matching keywords, so it holds up well against sarcasm, typos, and the general messiness of real survey data. 

It also supports coding in almost any language out of the box, no separate translation step needed, which makes it a strong fit for global research.

You get a human review step to check and adjust the output, full traceability back to the original responses, and flexible pricing that includes a pay-as-you-go option, so you can start small and grow into it rather than committing to a big contract upfront.

If you want to see how Blix handles your own data, you can try it on a real sample before deciding anything.

Key Takeaways

Start with your own needs. Then judge every tool against the same things: quality, speed, ease of use, control and flexibility, privacy and security, and cost. 

Test the platforms on your own data whenever you can, and look for a tool that fits naturally into the way your team already works.

Get those things right, and the right platform can save your team a huge amount of time, make open-ended analysis much easier, and help you get more value from the feedback you’re already collecting.

"

FAQ

What should I look for in verbatim coding software?

Prioritize accuracy, a real human-in-the-loop review process, traceability from each theme back to the original responses, and quantitative outputs that match how your team reports results.

How do I test a vendor before buying?

Ask to run a sample of your own real, unedited data, and compare the results against how your team would code the same responses by hand.

Is AI or human coding more accurate?

It depends on the tool. Older, keyword-based software often struggles with context, sarcasm, and nuance. More recent tools built on large language models look at the meaning of a response rather than the specific words used, which has closed much of that gap.

Gal Orian Harel
Co-Founder & CEO
Linkedin profile

Gal Orian Harel is the co-founder and CEO of Blix. With over 15 years of experience turning feedback into insights, products, and strategy, he’s passionate about helping research teams move faster and think smarter by transforming open-ended feedback into clear, actionable insights.

Still coding open ends manually? Save hours with Blix

Tired of manual coding? Talk to us

Save hours of manual work with AI powered open ends coding, with human-level quality and zero manual work.

Turn qualitative feedback into data and insights in minutes, with a few clicks.

Blix is trusted by top brands and market research firms worldwide:

Book a demo

You can reach us anytime via info@blix.ai

check icon

Thank you

We will contact you shortly to book a demo.
Oops! Something went wrong while submitting the form.

Please try again.