Skip to main content

A platform for describing, extracting, transforming, loading and serving open data.

Project description

Spinta is a framework to describe, extract and publish data (DEP Framework). It supports a great deal of data schemes and formats.

https://gitlab.com/atviriduomenys/spinta/badges/master/pipeline.svg https://gitlab.com/atviriduomenys/spinta/badges/master/coverage.svg

Purpose

  • Describe your data: You can automatically generate data structure description table (Manifest) from many different data sources.

  • Extract your data: Once you have your data structure in Manifest tables, you can extract data from multiple external data sources. Extracted data are validated and transformed using rules defined in Manifest table. Finally, data can be stored into internal database in order to provide fast and flexible access to data.

  • Publish your data: Once you have your data loaded into internal database, you can publish data using API. API is generated automatically using Manifest tables and provides extracted data in many different formats. For example if original data source was a CSV file, now you have a flexible API, that can talk JSON, RDF, SQL, CSV and other formats.

Features

  • Simple 15 column table format for describing data structures (you can use any spreadsheet software to manage metadata of your data)

  • Internal data storage with pluggable backends (PostgreSQL or Mongo)

  • Build-in async API server built on top of Starlette for data publishing

  • Simple web based data browser.

  • Convenient command-line interface

  • Public or restricted API access via OAuth protocol using build-in access management.

  • Simple DSL for querying, transforming and validating data.

  • Low memory consumption for data of any size

  • Support for many different data sources

  • Advanced data extraction even from dynamic API.

  • Compatible with DCAT and Frictionless Data Specifications.

Example

If you have an SQLite database:

$ sqlite3 sqlite.db <<EOF
CREATE TABLE COUNTRY (
    NAME TEXT
);
EOF

You can get a limited API and simple web based data browser with a single command:

$ spinta run -r sql sqlite:///sqlite.db

Then you can generate metadata table (I call it manifest) like this:

$ spinta inspect -r sql sqlite:///sqlite.db
d | r | b | m | property | type   | ref | source              | prepare | level | access | uri | title | description
dataset                  |        |     |                     |         |       |        |     |       |
  | sql                  | sql    |     | sqlite:///sqlite.db |         |       |        |     |       |
                         |        |     |                     |         |       |        |     |       |
  |   |   | Country      |        |     | COUNTRY             |         |       |        |     |       |
  |   |   |   | name     | string |     | NAME                |         | 3     | open   |     |       |

Generated data structure table can be saved into a CSV file:

$ spinta inspect -r sql sqlite:///sqlite.db -o manifest.csv

Missing peaces in metadata can be filled using any Spreadsheet software.

Once you done editing metadata, you can test it via web based data browser or API:

$ spinta run --mode external manifest.csv

Once you are satisfied with metadata, you can generate a new metadata table for publishing, removing all traces of original data source:

$ spinta copy --no-source --access open manifest.csv manifest-public.csv

Now you have matadata for publishing, but all things about original data source are gone. In order to publish data, you need to copy external data to internal data store. To do that, first you need to initialize internal data store:

$ spinta config add backend my_backend postgresql postgresql://localhost/db
$ spinta config add manifest my_manifest tabular manifest-public.csv
$ spinta migrate

Once internal database is initialized, you can push external data into it:

$ spinta push --access open manifest.csv

And now you can publish data via full featured API with a web based data browser:

$ spinta run

You can access your data like this:

$ http :8000/dataset/sql/Country
HTTP/1.1 200 OK
content-type: application/json

{
    "_data": [
        {
            "_type": "dataset/sql/Country",
            "_id": "abdd1245-bbf9-4085-9366-f11c0f737c1d",
            "_rev": "16dabe62-61e9-4549-a6bd-07cecfbc3508",
            "_txn": "792a5029-63c9-4c07-995c-cbc063aaac2c",
            "name": "Vilnius"
        }
    ]
}

$ http :8000/dataset/sql/Country/abdd1245-bbf9-4085-9366-f11c0f737c1d
HTTP/1.1 200 OK
content-type: application/json

{
    "_type": "dataset/sql/Country",
    "_id": "abdd1245-bbf9-4085-9366-f11c0f737c1d",
    "_rev": "16dabe62-61e9-4549-a6bd-07cecfbc3508",
    "_txn": "792a5029-63c9-4c07-995c-cbc063aaac2c",
    "name": "Vilnius"
}

$ http :8000/dataset/sql/Country/abdd1245-bbf9-4085-9366-f11c0f737c1d?select(name)
HTTP/1.1 200 OK
content-type: application/json

{
    "name": "Vilnius"
}

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

spinta-0.1.26.tar.gz (213.8 kB view details)

Uploaded Source

Built Distribution

spinta-0.1.26-py3-none-any.whl (312.6 kB view details)

Uploaded Python 3

File details

Details for the file spinta-0.1.26.tar.gz.

File metadata

  • Download URL: spinta-0.1.26.tar.gz
  • Upload date:
  • Size: 213.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.1.12 CPython/3.10.2 Linux/5.10.96-1-MANJARO

File hashes

Hashes for spinta-0.1.26.tar.gz
Algorithm Hash digest
SHA256 bdf69bcb228dbd62e88ca8f69105dab138a5ea70d0971403eaa48e4583df8e70
MD5 f4b4cd84b05f3c4107c543413c76ef4b
BLAKE2b-256 155bf2d07358b682f7bb5fa1fa872759fe21c89c6efdf9ab8616f004b37be57d

See more details on using hashes here.

File details

Details for the file spinta-0.1.26-py3-none-any.whl.

File metadata

  • Download URL: spinta-0.1.26-py3-none-any.whl
  • Upload date:
  • Size: 312.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.1.12 CPython/3.10.2 Linux/5.10.96-1-MANJARO

File hashes

Hashes for spinta-0.1.26-py3-none-any.whl
Algorithm Hash digest
SHA256 6f79d3909a00a4e6ef3d13b3ed6d95a3dfc70c121c57daee9a777aeb8be824de
MD5 b957d85d1cb56039d2bfe5f262b513c5
BLAKE2b-256 0ff7b509b656c163e37370698b3cbfcd4c700430f6e27c8b8f2ddf52376fb924

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page