Memory hook: Structure describes the data; workload describes how it is used.
Must remember
- Structured data has a defined tabular schema; semi-structured data such as JSON carries flexible structure; unstructured content includes images/audio and free text. CSV, JSON, XML, Parquet and Avro differ in schema, representation and analytical efficiency.
- A database organises data for supported access; a file store holds files/objects. Relational tables, key-value records, documents, graphs and column-family models serve different access patterns. Schema-on-write validates before storage; schema-on-read interprets data during use.
- OLTP handles frequent small transactions with consistency requirements; OLAP analyses large historical datasets. Batch processes bounded collections; streaming processes continuing events. Low arrival latency does not automatically guarantee exactly-once results.
- Data engineers build ingestion/transformation pipelines; database administrators manage database operation/security/performance; analysts model and interpret data for decisions. Responsibilities can overlap, but the role distinction helps select the right activity.
- Data quality includes validity, completeness, uniqueness, consistency and freshness. Governance covers ownership, access, lineage, retention and lawful use. A well-formatted record can still be factually wrong.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Checkout updates several related records | Transactional workload. |
| Analyse years of sales by product | Analytical workload. |
| Continuously evaluate sensor events | Streaming processing. |
Traps
- JSON is not necessarily unstructured.
- Real-time ingestion and real-time dashboards are separate stages.
- A file extension alone does not prove valid contents.
Active recall
1. Is a photograph structured tabular data?
No; it is normally unstructured content, though metadata may be structured.
2. Which role builds data pipelines?
A data engineer, typically.
3. What differs between OLTP and OLAP?
Transactional updates versus analytical exploration/aggregation of data.
4. Why use columnar formats for analytics?
They support efficient reading/compression of selected columns.
5. Does a successful ingestion prove quality?
No. Validate content, freshness and reconciliation rules.