Metadata-Version: 2.4
Name: rich_datatable
Version: 0.1.2
Summary: Typed pandas tables with metadata and data quality helpers.
Author-email: Scientith <dev@scientith.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/scienith/rich-datatable
Project-URL: Repository, https://github.com/scienith/rich-datatable
Project-URL: Issues, https://github.com/scienith/rich-datatable/issues
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas
Requires-Dist: loguru
Requires-Dist: pydantic
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov; extra == "dev"
Requires-Dist: black; extra == "dev"
Requires-Dist: isort; extra == "dev"
Requires-Dist: mypy; extra == "dev"
Dynamic: license-file

Rich DataTable (core)

Minimal, computation-only table library for:
- Loading CSV/Excel into a typed table
- Computing basic statistics
- Inferring column/table metadata
- Running data quality checks (QC) based on metadata

The stable public API is surfaced via `rich_datatable`. The legacy
`datatable_core` facade remains available for compatibility.

Install (editable)
- `pip install "rich-datatable @ git+https://github.com/scienith/rich-datatable.git@v0.1.2"` (Python >= 3.11)

Public API
- `DataTable`: minimal in-memory table with lazy helpers
- `MetadataInferenceThresholds`: optional inference configuration
- `SeverityText`: QC severity levels (for interpreting results)

Quickstart
1) Load and infer from a CSV/Excel file
```python
from rich_datatable import DataTable

dt = DataTable.load(
    "data.csv",
    table_name="my_table",
    sample_keys="id",  # set once here and reuse
)

df = dt.to_pandas()
meta = dt.get_metadata()
```

2) Adjust inference thresholds (optional)
```python
from rich_datatable import DataTable, MetadataInferenceThresholds

thresholds = MetadataInferenceThresholds(
    null_threshold=0.05,
    unique_threshold=0.95,
    categorical_threshold=0.75,
    allowed_value_threshold=0.5,
    mad_factor=4.5,
)

dt = DataTable.load("data.xlsx", thresholds=thresholds, sheet_name="Sheet1")
```

3) Run QC (data quality checks)
```python
# Prefer the convenience method (sample_keys stored in DataTable)
results, summaries, cell_errors, column_errors = dt.run_qc()
```

Design notes
- No I/O/report styling/visualization (e.g., Styled Excel, Plotly) is included.
- `rich_datatable` is the recommended import path for consumers.
- Internals (stats/inference/low-level enums/models) are kept private unless needed.

Public surface (kept minimal)
- `DataTable` — load + access data/metadata, run QC
- `MetadataInferenceThresholds` — inference configuration
- `SeverityText` — interpret QC severities

Non-goals
- Styled outputs, report generation, plotting, or persistence APIs are out of scope.

License
- MIT; see `LICENSE`.
