Back to blog
·9 min read

How to format CSV and JSON product catalogs for precise AI chatbot responses

How to format CSV and JSON product catalogs for precise AI chatbot responses

When e-commerce and B2B companies train their AI chatbot on product catalogs, many upload raw CSV or JSON exports straight from their ERP or PIM system. Without proper structure, the chatbot often swaps dimensions, reports incorrect prices, or misattributes accessories to the wrong parent item. This happens because database tables are designed for relational queries, not semantic language retrieval. This guide shows how to format CSV and JSON product catalogs to reduce attribute hallucinations and deliver far more reliable product specifications.

Why raw catalog exports cause attribute hallucinations

Traditional e-commerce databases and PIM systems store data in highly compressed or relational formats. A raw CSV export often uses cryptic column headers (such as "attr_12_val" or "qty_p3") that depend on a hierarchy defined elsewhere. When a modern RAG-based AI chatbot indexes this file, it breaks the dataset into text chunks for semantic retrieval.

If a chunk lands in the middle of a row-dense table without explicit column headers, the language model loses key-value context. The result is attribute hallucination: the chatbot guesses that a VESA mount specification belongs to a laptop on the next row, or assumes a numeric price includes tax when the catalog listed net pricing.

To ensure more reliable retrieval, every product record inside your catalog file must be self-contained and carry explicit context. Avoid three frequent structural pitfalls:

  • Deeply nested arrays:JSON trees where crucial attributes like "width" sit three levels deep under an anonymous key.
  • Cryptic column abbreviations:Headers that rely on internal legacy codes like "dim_w_mm" without plain-text semantics.
  • Missing product titles per row: CSV rows that list sub-SKUs without repeating the parent product name.

Best practices for CSV: column headers, flattening, and explicit context

CSV is a highly cost-effective format for chatbot knowledge bases because it is simple to update from export scripts. To ensure reliable AI retrieval, your CSV file must adhere to explicit context and flattened structure rules.

Make sure the header row uses descriptive, plain-language names. Instead of "Weight", use "Weight (kg)". Instead of "Price", write "Price incl. VAT (NOK)". Furthermore, every single row must include the full canonical product name, ensuring that if a chunk is retrieved in isolation, the AI model knows which product the attributes belong to.

Follow this structural checklist for CSV product files:

  • Flatten nested hierarchies: Avoid merged cells or spreading a single product over multiple rows. One row must equal one unique product or SKU.
  • Add a text summary column:Include a dedicated column named "Product Summary" that combines key attributes into natural sentences.
  • Maintain strict formatting: Use standard comma or semicolon delimiters, ensuring string fields containing internal punctuation are properly double-quoted.

Example of a flattened CSV row with explicit context:

Product Name,SKU,Weight (kg),Price incl. VAT (NOK),Product Summary
"Acme Monitor 27"" Pro",MON-27-PRO,5.2,4490,"27-inch display with VESA 100x100, weight 5.2 kg, price 4490 NOK incl. VAT."

Optimizing JSON: structure, key-value clarity, and text summaries

If your product catalog features variable specifications, complex compatibility matrices, or multi-tier options, JSON provides greater flexibility than CSV. However, poorly structured JSON can cause major retrieval gaps during vector search.

Avoid huge array structures that nest hundreds of items under a generic parent tag. Instead, represent your catalog as a list of independent, self-contained objects with semantic key-value pairs. For technical specs, prefer flat keys like "screen_size_inches": 27 over deep anonymous structures such as "specs": [{"id": 1, "v": "27"}].

A highly effective technique is adding a synthesized plain-text field inside each JSON object (e.g., "ai_summary"). This property joins the core specifications into a readable paragraph. During vector search, semantic embeddings hit this text summary with high confidence, while the LLM extracts exact values from the adjacent key-value parameters.

Before (hard for RAG) versus after (self-contained object):

// Before
{ "specs": [{ "id": 1, "v": "27" }] }

// After
{
  "product_name": "Acme Monitor 27 Pro",
  "sku": "MON-27-PRO",
  "screen_size_inches": 27,
  "vesa_mount_mm": "100x100",
  "weight_kg": 5.2,
  "price_incl_vat_nok": 4490,
  "ai_summary": "Acme Monitor 27 Pro is a 27-inch display with VESA 100x100 mounting, weighing 5.2 kg, priced at 4490 NOK incl. VAT."
}

Comparison of catalog formats for AI retrieval

Choosing between catalog formats depends on the complexity of your technical data and update frequency. The comparison table below outlines how format choices impact retrieval accuracy and hallucination risk.

FormatRetrieval accuracyHallucination riskRecommended use case
Raw ERP export (CSV)LowHighNot recommended for direct AI training without cleaning
Flattened and labeled CSVHighLowIdeal for large product catalogs and e-commerce
Semantic key-value JSONVery highMinimalBest choice for complex technical specifications

As the table shows, both flattened CSV and semantic JSON yield significantly lower error rates than raw database dumps. Investing time in a lightweight preprocessing script ensures your buyers receive trustworthy details and reduces costly order errors.

Practical QA and catalog deployment

Once your catalog file is properly formatted, perform a systematic quality assurance audit before going live. Test the assistant with queries that require precise attribute discrimination, such as comparing two adjacent SKUs or checking cross-compatibility between accessories.

Use a platform with source citations and RAG technology, so your team can verify exactly which text chunk or database row informed the response. If the assistant picks the wrong item row, refine your column headers or strengthen the text summary property.

With Chatly's platform, you can upload CSV and JSON data files seamlessly alongside web URLs and PDF documentation. The system parses structured datasets to preserve attribute context, transforming your catalog into a reliable 24/7 sales assistant.

Ready to turn your product catalog into a reliable AI assistant? Explore Chatly today or get in touch with our team to see how structured knowledge bases improve accuracy for your online store.