Every data project starts with a simple question: what shape is this data in? The answer decides where you can store it, which tools can read it and how much work it takes before you can analyze it. Data+ groups data into three shapes: structured, semi-structured and unstructured.
Structured data follows a fixed schema: every record has the same fields, each field has a defined type, and the data fits naturally into rows and columns. A table of orders in a relational database, with OrderID, CustomerID, OrderDate and Amount columns, is the classic example. Spreadsheets laid out as clean tables are structured too. Because the shape is known in advance, you can query structured data directly with SQL, and it is the easiest kind to aggregate, join and chart.
Semi-structured data carries its own labels but does not force every record into the same shape. JSON and XML are the common examples: each value sits next to a key or inside a tag that says what it is, but one record can have fields another lacks, and values can be nested (an order containing a list of line items). Log files with key=value pairs, email headers and many API responses are semi-structured. You usually parse or flatten this data into tables before analysis, and modern databases and tools such as pandas can read it directly.
Unstructured data has no data model that a query can use. Free text in emails, chat messages and documents, PDFs, images, audio and video all fall here. Most of the world's data is unstructured, and it often holds valuable information, such as the reason a customer is unhappy, but you need extra processing to turn it into analyzable fields. Examples include natural language processing to score sentiment, optical character recognition to pull text from scanned forms, or tags assigned by a person or a model.
Watch for mixed cases. An email is semi-structured in its headers (From, To, Date) and unstructured in its body. A CSV file is structured if every row has the same columns, even though it is just a text file. A database table can hold an unstructured column, such as a comments field. When a question asks you to classify data, look at whether a consistent schema describes it, whether it is self-describing with keys or tags, or whether it has no model at all.
Key terms
- Structured data
- Data that follows a fixed schema of fields and types, such as rows in a relational table.
- Semi-structured data
- Self-describing data with keys or tags but a flexible shape, such as JSON or XML.
- Unstructured data
- Data with no predefined model that queries can use, such as free text, images, audio and video.
- Schema
- The definition of the fields, types and relationships that data is expected to follow.
A support team wants to know why customers cancel. Account details and cancellation dates sit in a structured CRM table, ticket metadata arrives as JSON from the helpdesk API, and the real reasons are in free-text chat transcripts. The analyst joins the structured and JSON data on customer ID, then uses a text classification step to tag each transcript with a reason, turning unstructured text into a column that can be counted.
Check yourself
A web API returns records where some objects contain an optional 'discount' key and a nested list of items. How is this data classified?
Semi-structured. It is self-describing with keys, but records do not all share one fixed schema.
Why does unstructured data usually need extra processing before analysis?
It has no fields that queries can filter or aggregate, so you must first extract features such as sentiment, keywords or text from images.