Movatterモバイル変換


[0]ホーム

URL:


Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Sign up
Appearance settings

omniparser: a native Golang ETL streaming parser and transform library for CSV, JSON, XML, EDI, text, etc.

License

NotificationsYou must be signed in to change notification settings

jf-tech/omniparser

Repository files navigation

CIcodecovGo Report CardPkgGoDevMentioned in Awesome Go

Omniparser is a native Golang ETL parser that ingests input data of various formats (CSV, txt, fixed length/width,XML, EDI/X12/EDIFACT, JSON, and custom formats) in streaming fashion and transforms data into desired JSON outputbased on a schema written in JSON.

Min Golang Version: 1.16

Licenses and Sponsorship

Omniparser is publicly available underMIT License.Individual and corporate sponsorships are welcome and gratefullyappreciated, and will be listed in theSPONSORS page.Company-level sponsors get additional benefits and supportsgranted in theCOMPANY LICENSE.

Documentation

Docs:

References:

Examples:

In the example folders above you will find pairs of input files and their schema files. Then in the.snapshots sub directory, you'll find their corresponding output files.

Online Playground (not functioning)

UseThe Playground (may need to wait for a few seconds for instance to wake up)for trying out schemas and inputs, yours or existing samples, to see how ingestion and transform work.

As for now (2023/03/14), all of our previous free docker hosting solutions went away and we haven't found another one yet. For now please clone the repo and use./cli.sh as described in theGetting Started page.

Why

  • No good ETL transform/parser library exists in Golang.
  • Even looking into Java and other languages, choices aren't many and all have limitations:
    • Smooks is dead, plus its EDI parsing/transform is too heavyweight, needing code-gen.
    • BeanIO can't deal with EDI input.
    • Jolt can't deal with anything other than JSON input.
    • JSONata still only JSON -> JSON transform.
  • Many of the parsers/transforms don't support streaming read, loading entire input into memory - not acceptable in somesituations.

Requirements

  • Golang 1.16 or later.

Recent Major Feature Additions/Changes

  • 2024/06: v1.0.5 released:upgraded minimum go version to 1.16; enabled full ES6 feature support in javascript custom function.
  • 2022/09: v1.0.4 released: addedcsv2 file format that supersedes the originalcsv format with support of hierarchical and nested records.
  • 2022/09: v1.0.3 released: addedfixedlength2 file format that supersedes the originalfixed-length format with support of hierarchical and nested envelopes.
  • 1.0.0 Released!
  • AddedTransform.RawRecord() for caller of omniparser to access the raw ingested record.
  • Deprecatedcustom_parse in favor ofcustom_func (custom_parse is still usable forback-compatibility, it is just removed from all public docs and samples).
  • AddedNonValidatingReader EDI segment reader.
  • Added fixed-length file format support in omniv21 handler.
  • Added EDI file format support in omniv21 handler.
  • Major restructure/refactoring
    • Upgrade omni schema version toomni.2.1 due a number of incompatible schema changes:
      • 'result_type' ->'type'
      • 'ignore_error_and_return_empty_str ->'ignore_error'
      • 'keep_leading_trailing_space' ->'no_trim'
    • Changed how we handle custom functions: previously we always use strings as in param type as well as result paramtype. Not anymore, all types are supported for custom function in and out params.
    • Changed the way we package custom functions for extensions: previously we collected custom functions from allextensions and then passed all of them to the extension that is used; this feels weird, now only the customfunctions included in a particular extension are used in that extension.
    • Deprecated/removed most of the custom functions in favor of using 'javascript'.
    • A number of package renaming.
  • Added CSV file format support in omniv2 handler.
  • Introduced IDR node cache for allocation recycling.
  • IntroducedIDR for in-memory data representation.
  • Added trie based high performancetimes.SmartParse.
  • Command line interface (one-offtransform cmd or long-running httpserver mode).
  • javascript engine integration as a custom_func.
  • JSON stream parser.
  • Extensibility:
    • Ability to provide custom functions.
    • Ability to provide custom schema handler.
    • Ability to customize the built-in omniv2 schema handler's parsing code.
    • Ability to provide a new file format support to built-in omniv2 schema handler.

Footnotes


[8]ページ先頭

©2009-2025 Movatter.jp