Dataset Profiler

Describe each column of a CSV file with a guessed type, blanks, unique values and samples.

Files stay on your device No sign-up Free to use
How this works

The tool runs in this browser. Your file or text is not uploaded to UseFreeTools. Check this tool's limits for anything it may save on your device.

Privacy details

Profile the columns controls

Drop your CSV file here

or choose one from your device

One UTF-8 file up to 5 MiB, 10,000 rows and 100 columns.

    Paste the table to describe, or leave this empty and choose a file. The data is read and summarised; nothing is changed.

    Without a names row, columns are reported as column_1, column_2 and so on.

    On by default, so a cell holding only spaces counts as empty and " ABC " counts the same as "ABC".

    Sample values are the first few distinct non-empty values in file order.

    How to use Dataset Profiler

    1. Paste the table into the box, or leave it empty and choose a CSV or TSV file.
    2. Set the delimiter only when the detected one looks wrong.
    3. Select Profile the columns and read the per-column summary of type, empty rate and unique values.

    Example: Dataset Profiler

    Check a small order file before importing it, where one quantity is empty and one customer name repeats.

    You add
    Text: header id,name,qty then rows 0012,Ana,2, 0013,Bo, (an empty qty) and 0014,Ana,7.
    You get
    The result reports 3 rows, 3 columns and 1 of 9 cells empty. The id column is integer with 3 unique values and range 12 to 14, name is text with 2 unique values, and qty is integer with 1 empty and 2 unique values. The notes name id, and only id, as a column that can identify a row.

    Options

    Value types
    Each non-empty value gets the narrowest fitting type: integer, decimal, date, boolean or text. A column is only called numeric when every non-empty value is numeric; one text value makes the whole column mixed.
    Ignore spaces at the ends of a value
    On by default, so a cell holding only spaces counts as empty and a padded value counts the same as its trimmed form. Switching it off shows padding as a real difference in both the blank count and the unique count.
    Unique values
    The number of distinct non-empty values in a column. A column where that equals the row count has no repeated value, which is what makes it a candidate key; blank cells are never counted as a value.
    Range
    Shown only for a numeric column, as the smallest and largest number in it. The digits are read from the text, so 0012 appears in the samples as written but is compared as the number twelve for the range.

    Supported inputs and limits

    One UTF-8 CSV or TSV source, pasted or from a file, up to 5 MiB, 10,000 rows and 100 columns. Quoted commas, quotes and line breaks inside a cell are read as part of that cell, and an identifier such as 0012 keeps its leading zeroes in the samples and unique count. The table shows up to 100 columns and keeps up to five example values per column. This page reads and describes the data; it does not edit, clean or export it. The guessed types are a reading of the text, not a schema, so a date that arrives as 29 Sep 2026 reads as text and a numeric-looking identifier reads as a number.

    Where your input is processed

    This tool processes your input in this browser. Your text and files are not uploaded to UseFreeTools. Check this tool's limits for anything it may save on your device.

    Profiling before an import

    Most import failures trace back to a column whose real contents differ from its label: a numeric field with a stray marker, a key with duplicate values, or a date written in more than one shape. Reading the empty rate, unique count and a few sample values first turns those into a decision rather than an error halfway through a load.

    Why a guess is not a schema

    The type in this report comes from the characters in the sample only. A column that holds only digits this month can hold a prefix next month, and a date written in short form may arrive in another format in the next export. The report is evidence about the file in front of you, so confirm the intended type against the data source before writing a schema.

    Questions about Dataset Profiler

    Why is my date column shown as text?

    The type guess recognises only a few plain date shapes, such as 2026-09-29 and 29/09/2026, and keeps a written month as text because many formats are ambiguous. Check that column by hand before relying on it.

    Why is my ID column called integer when it holds codes?

    The guess describes the characters, not their meaning. A column of digits reads as a number even when the digits are a label. Treat the type as a hint and decide the real type for yourself.

    Why is a numeric column called mixed?

    At least one non-empty value is not numeric, such as n/a or a blank marker written as text. The column is reported as mixed so a numeric summary is never shown for data that partly is text.

    Does the profiler change my file or upload it?

    No. The table is read and summarised in this browser. Nothing is written, sent to a server, or changed on your device.

    What is a candidate key?

    A column with no empty cells and a different value in every row is named in the notes. It is a candidate only: a column can look unique in this sample and still repeat in the full dataset.

    Project manager: Tony Hines · Content updated 29 September 2026 · Report a problem