collate.html
来自「perl教程」· HTML 代码 · 共 752 行 · 第 1/3 页
HTML
752 行
<dd>
<p>-- see 3.2.2 Variable Weighting, UTS #10.</p>
</dd>
<dd>
<p>Makes the entry in the table completely ignorable;
i.e. as if the weights were zero at all level.</p>
</dd>
<dd>
<p>Through <a href="#item_ignorechar"><code>ignoreChar</code></a>, any character matching <a href="../../lib/Pod/perlfunc.html#item_qr_"><code>qr/$ignoreChar/</code></a>
will be ignored. Through <a href="#item_ignorename"><code>ignoreName</code></a>, any character whose name
(given in the <a href="#item_table"><code>table</code></a> file as a comment) matches <a href="../../lib/Pod/perlfunc.html#item_qr_"><code>qr/$ignoreName/</code></a>
will be ignored.</p>
</dd>
<dd>
<p>E.g. when 'a' and 'e' are ignorable,
'element' is equal to 'lament' (or 'lmnt').</p>
</dd>
</li>
<dt><strong><a name="item_katakana_before_hiragana">katakana_before_hiragana</a></strong>
<dd>
<p>-- see 7.3.1 Tertiary Weight Table, UTS #10.</p>
</dd>
<dd>
<p>By default, hiragana is before katakana.
If the parameter is made true, this is reversed.</p>
</dd>
<dd>
<p><strong>NOTE</strong>: This parameter simplemindedly assumes that any hiragana/katakana
distinctions must occur in level 3, and their weights at level 3 must be
same as those mentioned in 7.3.1, UTS #10.
If you define your collation elements which violate this requirement,
this parameter does not work validly.</p>
</dd>
</li>
<dt><strong><a name="item_level">level</a></strong>
<dd>
<p>-- see 4.3 Form Sort Key, UTS #10.</p>
</dd>
<dd>
<p>Set the maximum level.
Any higher levels than the specified one are ignored.</p>
</dd>
<dd>
<pre>
Level 1: alphabetic ordering
Level 2: diacritic ordering
Level 3: case ordering
Level 4: tie-breaking (e.g. in the case when variable is 'shifted')</pre>
</dd>
<dd>
<pre>
ex.level => 2,</pre>
</dd>
<dd>
<p>If omitted, the maximum is the 4th.</p>
</dd>
</li>
<dt><strong><a name="item_normalization">normalization</a></strong>
<dd>
<p>-- see 4.1 Normalize, UTS #10.</p>
</dd>
<dd>
<p>If specified, strings are normalized before preparation of sort keys
(the normalization is executed after preprocess).</p>
</dd>
<dd>
<p>A form name <code>Unicode::Normalize::normalize()</code> accepts will be applied
as <code>$normalization_form</code>.
Acceptable names include <code>'NFD'</code>, <code>'NFC'</code>, <code>'NFKD'</code>, and <code>'NFKC'</code>.
See <code>Unicode::Normalize::normalize()</code> for detail.
If omitted, <code>'NFD'</code> is used.</p>
</dd>
<dd>
<p><a href="#item_normalization"><code>normalization</code></a> is performed after <a href="#item_preprocess"><code>preprocess</code></a> (if defined).</p>
</dd>
<dd>
<p>Furthermore, special values, <a href="../../lib/Pod/perlfunc.html#item_undef"><code>undef</code></a> and <code>"prenormalized"</code>, can be used,
though they are not concerned with <code>Unicode::Normalize::normalize()</code>.</p>
</dd>
<dd>
<p>If <a href="../../lib/Pod/perlfunc.html#item_undef"><code>undef</code></a> (not a string <code>"undef"</code>) is passed explicitly
as the value for this key,
any normalization is not carried out (this may make tailoring easier
if any normalization is not desired). Under <code>(normalization => undef)</code>,
only contiguous contractions are resolved;
e.g. even if <code>A-ring</code> (and <code>A-ring-cedilla</code>) is ordered after <code>Z</code>,
<code>A-cedilla-ring</code> would be primary equal to <a href="../../lib/Pod/perlguts.html#item_a"><code>A</code></a>.
In this point,
<code>(normalization => undef, preprocess => sub { NFD(shift) })</code>
<strong>is not</strong> equivalent to <code>(normalization => 'NFD')</code>.</p>
</dd>
<dd>
<p>In the case of <code>(normalization => "prenormalized")</code>,
any normalization is not performed, but
non-contiguous contractions with combining characters are performed.
Therefore
<code>(normalization => 'prenormalized', preprocess => sub { NFD(shift) })</code>
<strong>is</strong> equivalent to <code>(normalization => 'NFD')</code>.
If source strings are finely prenormalized,
<code>(normalization => 'prenormalized')</code> may save time for normalization.</p>
</dd>
<dd>
<p>Except <code>(normalization => undef)</code>,
<strong>Unicode::Normalize</strong> is required (see also <strong>CAVEAT</strong>).</p>
</dd>
</li>
<dt><strong><a name="item_overridecjk">overrideCJK</a></strong>
<dd>
<p>-- see 7.1 Derived Collation Elements, UTS #10.</p>
</dd>
<dd>
<p>By default, CJK Unified Ideographs are ordered in Unicode codepoint order
but <code>CJK Unified Ideographs</code> (if <a href="#item_uca_version"><code>UCA_Version</code></a> is 8 to 11, its range is
<code>U+4E00..U+9FA5</code>; if <a href="#item_uca_version"><code>UCA_Version</code></a> is 14, its range is <code>U+4E00..U+9FBB</code>)
are lesser than <code>CJK Unified Ideographs Extension</code> (its range is
<code>U+3400..U+4DB5</code> and <code>U+20000..U+2A6D6</code>).</p>
</dd>
<dd>
<p>Through <a href="#item_overridecjk"><code>overrideCJK</code></a>, ordering of CJK Unified Ideographs can be overrided.</p>
</dd>
<dd>
<p>ex. CJK Unified Ideographs in the JIS code point order.</p>
</dd>
<dd>
<pre>
<span class="string">overrideCJK</span> <span class="operator">=></span> <span class="keyword">sub</span><span class="variable"> </span><span class="operator">{</span>
<span class="keyword">my</span> <span class="variable">$u</span> <span class="operator">=</span> <span class="keyword">shift</span><span class="operator">;</span> <span class="comment"># get a Unicode codepoint</span>
<span class="keyword">my</span> <span class="variable">$b</span> <span class="operator">=</span> <span class="keyword">pack</span><span class="operator">(</span><span class="string">'n'</span><span class="operator">,</span> <span class="variable">$u</span><span class="operator">);</span> <span class="comment"># to UTF-16BE</span>
<span class="keyword">my</span> <span class="variable">$s</span> <span class="operator">=</span> <span class="variable">your_unicode_to_sjis_converter</span><span class="operator">(</span><span class="variable">$b</span><span class="operator">);</span> <span class="comment"># convert</span>
<span class="keyword">my</span> <span class="variable">$n</span> <span class="operator">=</span> <span class="keyword">unpack</span><span class="operator">(</span><span class="string">'n'</span><span class="operator">,</span> <span class="variable">$s</span><span class="operator">);</span> <span class="comment"># convert sjis to short</span>
<span class="operator">[</span> <span class="variable">$n</span><span class="operator">,</span> <span class="number">0x20</span><span class="operator">,</span> <span class="number">0x2</span><span class="operator">,</span> <span class="variable">$u</span> <span class="operator">]</span><span class="operator">;</span> <span class="comment"># return the collation element</span>
<span class="operator">},</span>
</pre>
</dd>
<dd>
<p>ex. ignores all CJK Unified Ideographs.</p>
</dd>
<dd>
<pre>
<span class="string">overrideCJK</span> <span class="operator">=></span> <span class="keyword">sub</span><span class="variable"> </span><span class="operator">{()},</span> <span class="comment"># CODEREF returning empty list</span>
</pre>
</dd>
<dd>
<pre>
<span class="comment"># where ->eq("Pe\x{4E00}rl", "Perl") is true</span>
<span class="comment"># as U+4E00 is a CJK Unified Ideograph and to be ignorable.</span>
</pre>
</dd>
<dd>
<p>If <a href="../../lib/Pod/perlfunc.html#item_undef"><code>undef</code></a> is passed explicitly as the value for this key,
weights for CJK Unified Ideographs are treated as undefined.
But assignment of weight for CJK Unified Ideographs
in table or <a href="#item_entry"><code>entry</code></a> is still valid.</p>
</dd>
</li>
<dt><strong><a name="item_overridehangul">overrideHangul</a></strong>
<dd>
<p>-- see 7.1 Derived Collation Elements, UTS #10.</p>
</dd>
<dd>
<p>By default, Hangul Syllables are decomposed into Hangul Jamo,
even if <code>(normalization => undef)</code>.
But the mapping of Hangul Syllables may be overrided.</p>
</dd>
<dd>
<p>This parameter works like <a href="#item_overridecjk"><code>overrideCJK</code></a>, so see there for examples.</p>
</dd>
<dd>
<p>If you want to override the mapping of Hangul Syllables,
NFD, NFKD, and FCD are not appropriate,
since they will decompose Hangul Syllables before overriding.</p>
</dd>
<dd>
<p>If <a href="../../lib/Pod/perlfunc.html#item_undef"><code>undef</code></a> is passed explicitly as the value for this key,
weight for Hangul Syllables is treated as undefined
without decomposition into Hangul Jamo.
But definition of weight for Hangul Syllables
in table or <a href="#item_entry"><code>entry</code></a> is still valid.</p>
</dd>
</li>
<dt><strong><a name="item_preprocess">preprocess</a></strong>
<dd>
<p>-- see 5.1 Preprocessing, UTS #10.</p>
</dd>
<dd>
<p>If specified, the coderef is used to preprocess
before the formation of sort keys.</p>
</dd>
<dd>
<p>ex. dropping English articles, such as "a" or "the".
Then, "the pen" is before "a pencil".</p>
</dd>
<dd>
<pre>
<span class="string">preprocess</span> <span class="operator">=></span> <span class="keyword">sub</span><span class="variable"> </span><span class="operator">{</span>
<span class="keyword">my</span> <span class="variable">$str</span> <span class="operator">=</span> <span class="keyword">shift</span><span class="operator">;</span>
<span class="variable">$str</span> <span class="operator">=~</span> <span class="regex">s/\b(?:an?|the)\s+//gi</span><span class="operator">;</span>
<span class="keyword">return</span> <span class="variable">$str</span><span class="operator">;</span>
<span class="operator">},</span>
</pre>
</dd>
<dd>
<p><a href="#item_preprocess"><code>preprocess</code></a> is performed before <a href="#item_normalization"><code>normalization</code></a> (if defined).</p>
</dd>
</li>
<dt><strong><a name="item_rearrange">rearrange</a></strong>
<dd>
<p>-- see 3.1.3 Rearrangement, UTS #10.</p>
</dd>
<dd>
<p>Characters that are not coded in logical order and to be rearranged.
If <a href="#item_uca_version"><code>UCA_Version</code></a> is equal to or lesser than 11, default is:</p>
</dd>
<dd>
<pre>
rearrange => [ 0x0E40..0x0E44, 0x0EC0..0x0EC4 ],</pre>
</dd>
<dd>
<p>If you want to disallow any rearrangement, pass <a href="../../lib/Pod/perlfunc.html#item_undef"><code>undef</code></a> or <code>[]</code>
(a reference to empty list) as the value for this key.</p>
</dd>
<dd>
<p>If <a href="#item_uca_version"><code>UCA_Version</code></a> is equal to 14, default is <code>[]</code> (i.e. no rearrangement).</p>
</dd>
<dd>
<p><strong>According to the version 9 of UCA, this parameter shall not be used;
but it is not warned at present.</strong></p>
</dd>
</li>
<dt><strong><a name="item_table">table</a></strong>
<dd>
<p>-- see 3.2 Default Unicode Collation Element Table, UTS #10.</p>
</dd>
<dd>
<p>You can use another collation element table if desired.</p>
</dd>
<dd>
<p>The table file should locate in the <em>Unicode/Collate</em> directory
on <a href="../../lib/Pod/perlvar.html#item__inc"><code>@INC</code></a>. Say, if the filename is <em>Foo.txt</em>,
the table file is searched as <em>Unicode/Collate/Foo.txt</em> in <a href="../../lib/Pod/perlvar.html#item__inc"><code>@INC</code></a>.</p>
</dd>
<dd>
⌨️ 快捷键说明
复制代码Ctrl + C
搜索代码Ctrl + F
全屏模式F11
增大字号Ctrl + =
减小字号Ctrl + -
显示快捷键?