Omni Calculator logo

Credibility in AI and Tech

Can we really trust AI? What are the limits of an AI system’s ability to give correct answers? What happens when AI is wrong, and what is the impact on our daily lives?

AI can reduce the time and cost required to complete a task or project. However, it is important to understand its limitations and assess the trustworthiness of the growing range of AI platforms.

At Omni, we have always worked to show the impact of math in our users’ daily lives, as pointed out in the article “Omni Calculator brings math to the masses” published in TechCrunch. Fundamental subjects such as math underpin technological development, which is why our calculators are useful in many areas of life, including AI. Two examples of Omni Calculator’s contributions to discussions around AI are:

    ● AI Job Risk Calculator

    ● AI Water Footprint Calculator

We are not saying that you should not use AI, nor are we denying its relevance in your life. However, we believe we can help investigate the limits of current AI tools and their potential effects on problem-solving skills and behavior.

Being transparent about our own relationship with AI is part of this commitment. Read Omni’s Stance on AI to learn more about what AI means to us, and see our Editorial Policies for information on how we use AI in creating and reviewing our content.

So, if you are worried about these questions or would like to learn more about AI or related technologies, you are in the right place.

Explore our original research, reports, and articles on AI and technology, designed to help you evaluate these subjects critically.

UX Research Report: 1 in 2 Omni Calculator users see AI as usable for calculations, but not yet reliable

This report shows that confidence in an answer matters as much as interface clarity. Although overall satisfaction with chatbot interactions is high (72.8%), only 59.2% of users trust AI. Chatbot interfaces, therefore, still have room to improve when users need help with calculations.

86% of Engineers in the US Are Using AI. Only 6% Trust It Without Hesitation

We surveyed 402 engineers to find out why they’re using AI tools they don’t fully trust and how they’re making sure things don’t break in the process.

1 in 3 Americans Received Wrong Calculation Results Using AI. A study of 1014 Participants (2026)

A 2026 study of 1,014 Americans shows that 1 in 3 users get incorrect math results from AI. While 62% of people use AI for calculations, a significant trust gap exists between Gen Z and Boomers. Learn why AI “instability” leads to errors and how to use these tools reliably.

Gemini Scores Ahead in New ORCA Benchmark, Outpacing ChatGPT Accuracy

Gemini scores ahead in new ORCA benchmark, outpacing ChatGPT accuracy. Read this updated AI benchmark report to learn more.

Who Should Pay for AI’s Water Bill? A survey of 703 Americans (2026 study)

A survey of 703 Americans finds that AI’s hidden water cost worries the public: 43% say tech companies should be responsible. Gen Z is more aware than older generations, and nearly three-quarters would change their AI habits after learning how much water data centers consume.

The ORCA Benchmark Evaluates How Well AIs Deal with Everyday Math

Why can’t you trust the numbers from your favorite AI chatbot? We put five leading AI models to the test with 500 everyday math problems to uncover their surprising error rates and show you how to avoid costly mistakes.

Why AI Sounds Like an Expert and How to Make it Act Like One Too

Why does AI always give answers with certainty? Why does it never respond with “I don’t know”? In this article, we’re delving into the topic and proposing solutions to make AI chatbots more dependable.

1 in 3 Workers Report a Ghost Downsizing As AI Shrinks Their Teams — But Most Bosses Disagree

A dual survey of 665 American workers and 354 C-suite executives reveals four disconnects between what employees are experiencing on the ground and what their leaders are planning, communicating, and publicly acknowledging.

Is Claude really the best? We tested its capabilities against its competitors

Check out our 2026 ORCA V3 benchmark to see which logic powerhouse you should be using!

Which AI Chatbot Builds the Best Calculator?

AI is building calculators now, so we decided to see how well they actually perform.

Is AI Killing Your Workplace Personality?

A survey of 921 US workers reveals the AI personality shift reshaping how we write at work, and why Gen Z trusts AI messages least.

1 in 4 Workers Secretly Avoid AI at Work (2026 survey)

A 2026 survey of 918 US workers found 1 in 4 have secretly avoided AI at work without telling their boss. Here’s who does it and why.

Survey: Half of Workers Have Caught Their Bosses Using AI (2026)

Our 2026 survey of 913 US workers shows how common it is for managers to use AI, what they hand off, and why it costs trust and the personal touch.

Claude vs Gemini: Which Is Best for Finance Professionals?

We tested leading AIs’ capabilities in finance! Check out this article to see our findings.

Statistics in LLMs — Introduction to basic concepts

We present a journey that will guide you to the basic concepts behind the answers generated by the large language models for different prompts.

Statistics in LLMs — Clustered Standard Errors

Clustered Standard Errors is a journey where you will learn how to evaluate the performance of large language models with dependent samples.

Applying KL Divergence in LLM Quantization

Applying KL Divergence in LLM Quantization is a report that explains how to evaluate the performance of quantized large language models.

Here you can learn more about our experts who are helping reshape how we interpret AI and apply new technologies.​

Dawid Siuda

Dawid Siuda

Dawid holds a Bachelor’s in Finance and a Master’s in Business IT and has the corporate credentials of a financial services consultant at one of Europe’s major banks. He also tutors in AI and computer science, researches AI, and evaluates how large language models handle complex computational reasoning.

Learn more about Dawid

Reyhaneh Mansouri

Reyhaneh Mansouri, PhD

Reyhaneh is a research writer and digital PR specialist at Omni Calculator, where she turns data into stories that help people and journalists. Her articles were featured in several media outlets, including Forbes, Daily Mail, Fox Business, The Hill, CNBC-Make It, Newsweek, and The Business Journal.​

Learn more about Reyhaneh

Sam Balboa

Sam Balboa

Sam is a Partnerships Project Manager and Digital PR Specialist at Omni Calculator. She turns complex data into stories people want to read, drawing on her BFA in Information Design and MA in Innovation Management. She also contributes to digital PR reports, press outreach, and link-building initiatives.

Learn more about Sam

Anna Kołota

Anna Kołota

Anna is a product designer with more than six years of experience. She combines her background in business, cognitive science, and service design to create meaningful user experiences. She is also an expert in exploring human-robot interactions and virtual reality, blending analytical thinking with creative problem-solving.

Learn more about Anna

Anna Szczepanek

Anna Szczepanek, PhD

Anna is a mathematician at the Jagiellonian University in Kraków. Bridging the gap between abstract numbers and the physical world, she focuses her academic research on mathematical physics and applied mathematics. Within the field of AI, her work has touched upon the mathematical foundations of modeling influence in social networks.

Learn more about Anna

Joanna Smietanska-Nowak

Joanna Śmietańska-Nowak, PhD

Joanna is a physics lecturer at AGH University of Kraków. Her diverse research experience spans protein crystallography, photocatalysis, and advanced energy storage systems, including supercapacitors and lithium-ion batteries. She has hands-on experience in neural network design, predictive modeling, and model evaluation, with training in LLMs and prompt engineering for academic applications.

Learn more about Joanna

Claudia Herambourg

Claudia Herambourg

Claudia holds a Bachelor’s degree in English literature and mathematical studies from University College Cork, Ireland. She is an aspiring computational linguist with a passion for both words and numbers. She is also working to unravel the mathematical laws governing semantic change using large language models.

Learn more about Claudia

João Rafael Lucio dos Santos

João Rafael Lucio dos Santos, PhD

João earned his doctorate from the University of Rochester. Now a university professor in Brazil, he brings extensive expertise in field theory, cosmology, and astrophysics to his science communication efforts. He has also applied his analytical expertise in the AI sector as a Prompt Writer, specializing in guiding and evaluating advanced large language models.

Learn more about João

Julia Kopczyńska

Julia Kopczyńska

Julia is a PhD student in microbiology at the Polish Academy of Sciences. She is a microbiologist who takes a holistic approach to health. She also worked in laboratories in Spain and the Netherlands, where she deepened her microbiological knowledge and research skills. She is currently exploring the application of AI and data analysis for biomedical research.

Learn more about Julia

Wojciech Sas

Wojciech Sas, PhD

Wojciech holds a doctorate from the Institute of Nuclear Physics PAN in Kraków. He currently works as a physicist in Zagreb, investigating how various materials behave under extreme conditions like low temperatures and high pressures. His work combines advanced experimentation with AI-driven data analysis to uncover complex patterns in material behavior.

Learn more about Wojciech

Featured by media

Here you can find our AI and technology articles, reports, and calculators from Omni Calculator that have been featured in the tech media:

Forbes logo

AI Is Changing How You Communicate. Here’s How To Stay Authentic. Excerpt: “Many people are using AI to communicate including about 22% of individual contributors and 52% of senior leaders, according to data from Omni Calculator.”

Entrepreneur logo

Before You Add AI to Your Product, Ask These 7 Questions. Excerpt: “In fact, in the third iteration of our ORCA Benchmark (Omni Research on Calculation in AI), which tests leading AI models on math and logic tasks, the best-performing free model (Grok 4.20) scored 70.4% accuracy, while Claude and ChatGPT came in at 53.2% and 48.4%, respectively.”

Inc logo

Forget Layoffs: This Quiet Workplace Tactic Is Shucking Headcounts Without Anyone Noticing. Excerpt: “In the past, however, relief usually arrived when the financial pressures that led to reductions abated, allowing recruitment to resume. But that cycle may now be ending. A new survey by Polish tech research firm Omni Calculator found nearly a third of employees reporting “ghost downsizing” to be on the rise, often among employers automating tasks of newly vacated positions with AI.“

The Globe and Mail logo

This AI Stock Is Down 13%, but Could Be the Safest One Out There. Excerpt: “A recent Omni Calculator study put Claude Sonnet 5 and Gemini 3.6 Flash through real finance tasks. The models had to build interactive calculators, solve complicated math problems, and answer financial-knowledge questions, including ones involving recent tax changes.”

euronews logo

Which AI chatbot is the best at simple math? Gemini, ChatGPT, Grok put to the test. Excerpt: “A recent study advises caution. The Omni Research on Calculation in AI (ORCA) shows that when you ask an AI chatbot to perform everyday math, there is roughly a 40 per cent chance it will get the answer wrong. Accuracy varies significantly across AI companies and across different types of mathematical tasks.”

The Register logo

AI is actually bad at math, ORCA shows. Excerpt: “According to their study, distributed via preprint service arXiv and on Omni Calculator’s website, ChatGPT-5, Gemini 2.5 Flash, Claude Sonnet 4.5, Grok 4, DeepSeek V3.2 achieved only 45–63 percent accuracy, with errors mainly related to rounding (35 percent) and calculation mistakes (33 percent).“

Techradar logo

Everyone’s switching from ChatGPT to Claude — but new tests say neither is the smartest free AI, and the real winner might surprise you. Excerpt: “ChatGPT is still the most popular AI chatbot around, even with the exodus that's underway to Claude, but is it the cleverest? A new report from OmniCalculator suggests that ChatGPT might not be the smartest AI around.“

The Hill logo

Don’t throw the generative baby out with the AI bathwater. Excerpt: ”The team behind Omni Calculator created the ORCA Benchmark to test this risk and found that no leading model scored above 63 percent on real-world calculation tasks. And generative AI is appearing in more government workflows and business systems every month, which means the cost of misreading its strengths and weaknesses keeps rising.”

R&D World logo

How developers and engineers are learning to work with AI they don’t fully trust. Excerpt: “An Omni Calculator survey of 403 U.S.-based engineers and engineering students, conducted in January 2026, reveals a similar landscape in engineering disciplines.”

Yahoo Finance logo

Omni Calculator Publishes ORCA V3 Research Report on AI Model Performance in Quantitative Reasoning. Excerpt: “Omni Calculator stated that the ORCA initiative was developed to provide additional transparency into AI model performance in mathematical and logical reasoning tasks and to support evaluation methods focused on real-world quantitative use cases. The full ORCA V3 report, titled Is Claude Really the Best?, is available on the Omni Calculator website.“

datacamp logo

29 Top Data Scientist Interview Questions For All Levels. Excerpt: “The confidence interval is a range of estimates for an unknown parameter that you expect to fall between a certain percentage of the time when you run the experiment again or similarly re-sample the population.“

Learn G2 logo

Artificial Neural Network: Applications and Software in 2024. Excerpt: “This function has a mathematical end result in the shape of an ‘S’ curve and is used when probabilities are the key criteria to determine whether the neuron should be activated. So, at any point, you can calculate the slope of this curve. The value of this function lies between 0 and 1.”

CEO World Magazine logo

Should You Trust AI with Your Numbers?. Excerpt: “This is where the promise of AI assisted analysis collides with a harder truth: if the arithmetic is untrustworthy, the story becomes unsafe to act on. The team behind Omni Calculator built the ORCA Benchmark to test that risk in everyday math, and no leading model scored above 63 percent on real-world tasks.“

Rebellion Research logo

AI-Ready Regions And Teams Are Redefining Tech Strategy. Excerpt: “A mechanical engineer pulls up a chat window, drops in a set of unit conversions, and watches a clean answer appear in seconds. Ten minutes later, the same engineer reaches for a scratch pad, runs the math again, and cross-checks the result against a familiar formula. That pattern has become the default rhythm of modern engineering work, and the AI adoption report from Omni Calculator puts hard numbers behind it.“

Tidio logo

ChatGPT for Customer Service: Best OpenAI Use Cases. Excerpt: “Users themselves flag the UX gap: research on AI chatbot interfaces found that top frustrations include errors or uncertainty in answers, overly long responses, and lack of visible reasoning – issues that compound trust problems in customer service contexts.“

Digital Journal logo

As AI writes more of our messages, are we losing our workplace personalities?. Excerpt: “A recently published survey from Omni Calculator, based on responses from office workers, suggests that many employees are beginning to notice a subtle but important shift. While AI may improve grammar, eliminate typos, and streamline communication, it may also be erasing some of the human characteristics that make workplace relationships meaningful.”

Techloy logo

Data Integrity is the Real Frontier of Modern Tech. Excerpt: “This is where an “error budget” comes in. By using a percent error calculator to monitor your system’s output, you can actually quantify how much trust you’re losing. If a ride-hailing algorithm predicts a 5-minute arrival but it consistently takes 7, that 40% error isn’t a glitch, it’s a broken promise.”

PR Newswire logo

ORCA Benchmark Reveals How AI’s Core Design Makes It Unreliable for Everyday Math. Excerpt: “Omni Calculator today released the findings of the ORCA (Omni Research on Calculation in AI) Benchmark, a comprehensive study evaluating leading AI chatbots on everyday math. The results are stark: users have a significant chance of receiving a wrong answer for calculable tasks, ranging from splitting a bill to projecting investment returns.”

Einpresswire logo

Gemini 3 Flash Crushes ChatGPT-5.2 in Accuracy Test - ORCA Benchmark Update. Excerpt: “The results are in for the second ORCA (Omni Research on Calculation in AI) Benchmark, and the leaderboard looks very different than it did two months ago. Gemini 3 Flash has surged to the top, becoming the first model to solve nearly three-quarters of real-world math and logic problems correctly.“

PRWeb logo

Omni Calculator Reveals Why AI Struggles With Precision and Trust in Calculations. Excerpt: “This benchmark will measure how accurately AI models, such as ChatGPT 5, Gemini 2.5 Flash, Claude Sonnet 4.5, and DeepSeek V3.2, solve 500 real-world, everyday calculation prompts–the same verified problems Omni Calculator handles daily. When AI Sounds Like an Expert, How to Make It Act Like One Too. Large language models (LLMs) are designed to predict text patterns, not to compute verified answers.“

Cited in technical papers

Here you can find our calculators cited in tech peer-reviewed journals and academic publications:

SIAM logo

Society for Industrial and Applied Mathematics

Article “Variationally Consistent Hamiltonian Model Reduction” cited the Speed of Sound in Solids Calculator.

Energy Technology logo

Energy Technology, Wiley

Article “Impact of the Cu Current Collector Grade and Ink Solid Fraction on the Electrochemical Performance of Silicon-Based Electrodes for Li-Ion Batteries” cited the Vickers Hardness Number Calculator.

NETECH logo

Nuclear Engineering and Technology, Elsevier

Article ”Empirical estimation of human error probabilities based on the complexity of proceduralized tasks in an analog environment” cited the Critical Value Calculator.

Cleaner Engineering and Technology logo

Cleaner Engineering and Technology, Elsevier

Article “Indoor CO2 direct air capture and utilization: Key strategies towards carbon neutrality” cited the Heat Loss Calculator.

MIJST logo

MIST International Journal of Science and Technology, MIST

Article “Link Budget Analysis in Designing a Web-application Tool for Military X-Band Satellite Communication” cited the Azimuth Calculator.

IJIM logo

International Journal of Interactive Mobile Technologies, OJS

Article “Sensor Based Algorithm for Self-Navigating Robot Using Internet of Things (IoT)” cited the Azimuth Calculator.

Agriengineering logo

Agriengineering, MDPI

Article “Development of a Small Dual-Chamber Solar PV-Powered Evaporative Cooling System for Fruit and Vegetable Cooling with Techno-Economic Assessment” cited the NPV Calculator.

Microplastics logo

Microplastics, MDPI

Article “An Open-Source Computer-Vision-Based Method for Spherical Microplastic Settling Velocity Calculation” cited the Water Density Calculator.

Applied Sciences logo

Applied Sciences, MDPI

Article “Development of a Prototype Solution for Reducing Soup Waste in an Institutional Canteen” cited the Truncated Cone Volume Calculator.