Skip to main content

Unicode aware lexicon for ZCTextIndex

Project description

Motivation

The standard ZCTextIndex lexicon only deals with 8-bit strings (and only if you get the zope.conf locale setting right). It does not handle Unicode or UTF-8. UnicodeLexicon fills this gap.

Installation

This product adds a ZCTextIndex Unicode Lexicon type to Zope. The lexicon comes with word splitters, stop word removers, a case normalizer, and two accent normalizers.

If you have GenericSetup installed, you can use the included extension profile to create a UnicodeLexicon in your portal_catalog and update the Title, Description, and SearchableText ZCTextIndexes.

There is no upgrade path from UnicodeLexicon 1.0. If you have 1.0 on your system, you have to delete and recreate the lexicon.

Pipeline Elements

The splitter works with all languages that separate words with whitespace characters.

The stop word remover knows about English language stop words only.

The accent normalizer comes in two flavors. There is a normalizer for Latin and Western European text (fr, es, pt, it, en, nl), and one for German and Scandinavian text (de, dk, no, se, fi, is). The latter keeps the umlaut characters ä, ö, and ü in tact.

Custom Pipeline Elements

Additional pipeline elements can be registered via ZCML. E.g.:

<configure
  xmlns="http://namespaces.zope.org/zope"
  xmlns:unicodelexicon="http://namespaces.zope.org/unicodelexicon">

  <include package="Products.UnicodeLexicon" file="meta.zcml" />

  <unicodelexicon:registerPipelineElement
    group="Accent Normalizer"
    name="Normalize accented chars (Custom text)"
    factory="my.package.pipeline.MyCustomNormalizer"
    />

</configure>

Default Encoding

The lexicon assumes either Unicode or UTF-8. If your application uses a different encoding, you can override the default by registering the encoding as a utility:

<configure
  xmlns="http://namespaces.zope.org/zope">

  <utility
    provides="Products.UnicodeLexicon.interfaces.IDefaultEncoding"
    component="my.package.pipeline.defaultEncoding"
    />

</configure>

Changelog

2.2 - 2011-01-30

  • Allow to override the default encoding in ZCML. [stefan]

2.1 - 2011-01-26

  • Add ability to register pipeline elements in ZCML. [stefan]

  • Fix a bug when updating PipelineFactory. [stefan]

2.0 - 2011-01-21

  • Add an ordered PipelineFactory. [stefan]

  • Add an accent-normalizing pipeline element originally contributed by Marc-Auréle Darche. [stefan]

  • Release as Python egg. [stefan]

1.0 - 2006-08-14

  • Initial release. [stefan]

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

Products.UnicodeLexicon-2.2.zip (26.5 kB view details)

Uploaded Source

File details

Details for the file Products.UnicodeLexicon-2.2.zip.

File metadata

File hashes

Hashes for Products.UnicodeLexicon-2.2.zip
Algorithm Hash digest
SHA256 98488a725a281679674c0fa80bf9a160b1a25a29207a669df115de8c1e311ed5
MD5 67cfba7757f9a7b14c4d23ee43ef4c0b
BLAKE2b-256 4d11e2fa879ab94baa7aa01bb6dd7b28de78f101f6679f8907e6abc25cf6eb04

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page