Comparing 13 Chat AI Products Across 8 Metrics: Pitfalls for Business Use
Comparing 13 conversational AIs with 8 business-focused metrics. Covers deployment types, free-plan terms traps, and IP indemnity handling.
According to a report by songchong on Qiita, a business-oriented comparison of 13 major conversational AI products was published on September 3, 2026. It examines primary sources such as terms of service, safety white papers, and pricing tables from August to September 2026. It covers products frequently mentioned in the field, including ChatGPT, Claude, Gemini, and Microsoft 365 Copilot. It aims to answer questions about safety, business fit, effectiveness, and manageability faced by executives and managers. It can also be read as a warning about the risks of using personal use for business purposes.
“Is it really safe to use ChatGPT at work in terms of data leaks and copyright?” “We hear about Microsoft 365 Copilot and Gemini too, but which one is actually useful for our work?”
A distinctive feature is that these questions are the starting point. The comparison does not compete on levels of intelligence. It aims to provide decision-making material in light of the confidentiality to be protected, budgets, and frontline operational capabilities. It is also important that it explicitly states it is not a ranking recommending specific products.
Overview of the 13 Products Compared from a
Business Perspective
This comparison is positioned as part of a series. It drills down into 13 conversational AI products from an overview that classified 66 business automation products into 6 roles. Conversational AI is defined as a tool for operating large language models in conversational form. It is distinguished from AI agents that autonomously operate business systems. It is described as an advisor that you consult on screen, returning text, summaries, ideas, and code.
The breakdown of the 13 products is diverse. There are 9 general web-based types, 2 office-software-embedded types, and 2 on-premise setup types. The general web-based types are ChatGPT, Claude, Gemini, Perplexity, Genspark, Grok, DeepSeek, Felo, and exaBase Generative AI. The office-software-embedded types are Microsoft 365 Copilot and Gemini Notebook. Gemini Notebook corresponds to the former NotebookLM. The on-premise setup types are Ollama and tsuzumi.
Evaluation is conducted using 8 original indicators. It is designed to address the 4 major questions that arise in adoption decisions. The 4 major questions are safety, business fit, effectiveness, and manageability. Vendors’ marketing claims vary in granularity and are difficult to compare directly. The significance therefore lies in independently defining yardsticks for cross-comparison.
Ratings are in four levels: ◎, ○, △, and —. They are organized based on each company’s published official terms, specifications, and pricing structures. They are presented as overall evaluations from the perspective of business use by SMEs. Supporting clauses and sources are listed in Chapters 2 and 6 of the original text. Legal protections such as copyright are not scored but handled separately.
Differences in Deployment Models That
Determine Ease of Adoption
The first hurdle to adoption is ease of getting started. Learning burden and technical requirements are key. Deployment models are organized into 4 categories. Operational burden differs greatly by category.
The first is the browser-based general web type. Nine products fall into this category and can be used immediately after registration. An advantage is that it does not depend on device performance. It is also a form in which employees often start using it on their own initiative. It has the downside that administrative departments struggle to grasp actual usage.
The second is the browser-based office-software-embedded type. Microsoft 365 Copilot and Gemini Notebook fall into this category. They blend into Word, Excel, Teams, and Google Drive environments. The strength is the low need to learn new operations. Existing permission management and storage infrastructure may be reused.
The third is the on-premise setup type for local execution on devices. One product, Ollama, falls into this category. It is installed and run on local computers or in-house servers. It can greatly reduce transmission of information to external clouds. On the other hand, high-performance GPUs and large-capacity storage are essential. Operation depends on the skills of the IT department.
The fourth is the on-premise setup type as a dedicated company infrastructure. One product, NTT’s tsuzumi, falls into this category. It is built in a dedicated partition such as Microsoft Azure or in on-premise facilities. It targets organizations with strict confidentiality requirements, such as financial institutions. It requires individual quotations and implementation effort. The criterion is ease of starting corporate use.
The perspective of stable operation is also essential. The greater the dependence on conversational AI, the greater the impact during outages. The simultaneous outage seen in Unexplained Simultaneous Outage of ChatGPT, Claude, and Grok offers an operational lesson. Combining embedded and web types and preparing fallback procedures will be challenges. Evaluation of continuity, not just ease of adoption, is required.
Terms-of-Service Traps Lurking in Free Plans
for Commercial Use
The existence of free plans accelerates adoption. However, continued business use of free plans may violate the terms. This survey clearly points out this issue. Clauses from 3 products are cited as specific examples.
Perplexity limits use to personal non-commercial use under its consumer terms. Business use falls under commercial purposes and violates prohibited acts. Felo prohibits applying as personal use despite corporate or commercial-purpose use. Grok stipulates that functional use for evaluation purposes is limited to personal non-commercial use. All of these illustrate the boundary between free use and business use.
This issue is directly linked to data-leak prevention. Use with personal accounts is beyond record retention and auditing. There is a risk that former employees’ accounts remain and chat histories stay outside the company. Whether data is used for training and deletion procedures also differ by product. Switching to corporate contracts can be seen as a prerequisite for management control.
In SMEs, tacitly tolerated use tends to spread. This is because convenience comes first and terms checks are postponed. Taking stock of actual usage and codifying internal rules is urgent. It is necessary to check not only whether free trials are allowed but also the licensing conditions for commercial use. Billing structures and usage limits should also be understood at the same time.
Limits of Feature Comparisons and Clarifying
Practical Boundaries
The scope of features affects usefulness in the field. However, vendors’ feature lists are inconsistently defined. A simple presence-or-absence table cannot be used for practical decisions. The concept of a practical boundary was therefore introduced.
The boundary is an idea to clarify the division of roles among tools. Conversational AI is positioned as a tool that gives answers but does not take action. It is clearly distinguished from entities that autonomously operate systems. While suited for consultation, summarization, and ideation support, execution is left to separate platforms. This clarification is seen as helpful in curbing excessive expectations.
The latest features as of 2026 are also subject to comparison. Document creation assistance, spreadsheet integration, and search integration are advancing across products. However, even with similar names, operating scope and permission management differ. Office-software-embedded types are strong in contextual reference. General web types are strong in cross-cutting research and organizing ideas.
In selection, fit with operations must be verified concretely. It is necessary to distinguish whether the use is drafting standard documents, summarizing minutes, or creating code. Cross-checking against the confidentiality classification of information handled is also essential. Prohibited inputs and handling of outputs should be defined at the trial stage. Alignment with operational procedures is seen as more important than sophistication of features.
Evaluation Decision to Treat IP Indemnity
Separately
IP indemnity was excluded from the scoring indicators and treated separately. This is because the conditions for legal protection are complex and unsuitable for scoring. This decision can be judged as practically reasonable. Understanding the applicable conditions matters more than the mere presence of indemnity.
Indemnity is usually attached to paid or corporate plans. Scope, procedures, and exclusions vary by provider. Originality of outputs and relations to third-party rights are also not uniform. In many cases, compliance with procedures by the user is a prerequisite. It is an area to be confirmed individually as a matter of fact.
In the field, ways of using outputs are diversifying. Use in externally published materials, proposals, and embedding in code is increasing. Confusion will arise unless assignment of responsibility for rights handling is defined. Operations for record retention and source verification are required. Coordination between legal departments and frontline teams is seen as key.
Perspectives for SMEs Deciding Whether to Adopt
Selection by SMEs differs from that by large enterprises. IT staffing and budgets are limited. Balancing ease of learning with management control is a challenge. The 8 indicators can be evaluated as designed to address this reality.
Office-software-embedded types impose a small learning burden. They have the advantage of using existing authentication infrastructure and storage locations. On the other hand, dependence on a specific platform deepens. They are susceptible to changes in pricing structures and usage limits. Procedures for reviewing contract terms need to be secured.
On-premise setup types offer major advantages in control. They enable reduced external transmission and in-house retention of records. However, implementation effort and hardware requirements are heavy. Securing personnel to handle operations is a prerequisite. Limiting application to operations with high confidentiality classifications is seen as realistic.
General web types are suited for ideation support and research. Adoption is fast, but management control becomes a challenge. Standardization on corporate plans and thorough enforcement of usage guidelines are essential. Choices must be made in light of the company’s confidentiality level and budget. The comparison table is seen as usable as a basis for internal approval documents.
Editorial Opinion
In the short term, responses to commercial-use restrictions on free plans are expected to advance. Development of internal rules and switching to paid plans may accelerate. Stocktaking of actual usage by administrative departments is expected to spread. Repurposing comparison tables for approval documents is also expected to increase.
In the long term, differentiated use of UI-integrated embedded types and on-premise setup types is expected to take hold. Highly confidential operations may remain in dedicated company environments, while routine consultations may be consolidated into external services. Selection criteria are expected to shift from sophistication of intelligence to operational control. The role of procurement departments is expected to expand.
The remaining issue is whether frontline teams can understand the applicable conditions for IP indemnity. It is necessary to verify whether the wording of terms matches actual operations. Responsibility for managing former employees and record retention is also seen as unclear. Disclosure of actual operational practices is evaluated as the next challenge.
References
- “ChatGPT・Copilot・Gemini… 業務視点でどう選べるか? ChatAI 13 製品を「8つの独自指標」で比較してみた|中小企業・管理職のための導入判断ガイド”, by songchong — Qiita, 2026-09-03T15:10:50.000Z (ARR)
- Source URL: https://qiita.com/songchong/items/78a9d9c28506ff90d331?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items
Frequently Asked Questions
- What are the 8 original indicators used in the comparison of the 13 products?
- They are yardsticks to answer the 4 major questions of safety, business fit, effectiveness, and manageability. They include ease of adoption and functional performance, organized as ◎, ○, △, — based on official terms and specifications. IP indemnity is not scored but handled separately.
- What should be noted when using free conversational AI for business?
- Perplexity, Felo, Grok and others have clauses limiting use to personal non-commercial use. Business use may violate the terms. Use with personal accounts is beyond audit and history management, so switching to corporate plans and establishing internal rules is necessary.
- Which deployment model should SMEs choose?
- Office-software-embedded types are practical to minimize learning burden. For highly confidential operations, on-premise setup types such as Ollama and tsuzumi are worth considering. General web types suit research and organizing ideas, but standardization on corporate plans and classification management of input information are prerequisites.
Comments