readme

来自「The GNU MP Bignum Library」· 代码 · 共 114 行

TXT

114 行

Copyright 2001 Free Software Foundation, Inc.This file is part of the GNU MP Library.The GNU MP Library is free software; you can redistribute it and/or modifyit under the terms of the GNU Lesser General Public License as published bythe Free Software Foundation; either version 3 of the License, or (at youroption) any later version.The GNU MP Library is distributed in the hope that it will be useful, butWITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITYor FITNESS FOR A PARTICULAR PURPOSE.  See the GNU Lesser General PublicLicense for more details.You should have received a copy of the GNU Lesser General Public Licensealong with the GNU MP Library.  If not, see http://www.gnu.org/licenses/.                   INTEL PENTIUM-4 MPN SUBROUTINESThis directory contains mpn functions optimized for Intel Pentium-4.The mmx subdirectory has routines using MMX instructions, the sse2subdirectory has routines using SSE2 instructions.  All P4s have these, theseparate directories are just so configure can omit that code if theassembler doesn't support it.STATUS                                cycles/limb	mpn_add_n/sub_n            4 normal, 6 in-place	mpn_mul_1                  4 normal, 6 in-place	mpn_addmul_1               6	mpn_submul_1               7	mpn_mul_basecase           6 cycles/crossproduct (approx)	mpn_sqr_basecase           3.5 cycles/crossproduct (approx)                                   or 7.0 cycles/triangleproduct (approx)	mpn_l/rshift               1.75The shifts ought to be able to go at 1.5 c/l, but not much effort has beenapplied to them yet.In-place operations, and all addmul, submul, mul_basecase and sqr_basecasecalls, suffer from pipeline anomalies associated with write combining andmovd reads and writes to the same or nearby locations.  The movqinstructions do not trigger the same hardware problems.  Unfortunately,using movq and splitting/combining seems to require too many extrainstructions to help.  Perhaps future chip steppings will be better.NOTESThe Pentium-4 pipeline "Netburst", provides for quite a number of surprises.Many traditional x86 instructions run very slowly, requiring use ofalterative instructions for acceptable performance.adcl and sbbl are quite slow at 8 cycles for reg->reg.  paddq of 32-bitswithin a 64-bit mmx register seems better, though the combinationpaddq/psrlq when propagating a carry is still a 4 cycle latency.incl and decl should be avoided, instead use add $1 and sub $1.  Apparentlythe carry flag is not separately renamed, so incl and decl depend on allprevious flags-setting instructions.shll and shrl have a 4 cycle latency, or 8 times the latency of the fastestinteger instructions (addl, subl, orl, andl, and some more).  shldl andshrdl seem to have 13 and 15 cycles latency, respectively.  Bizarre.movq mmx -> mmx does have 6 cycle latency, as noted in the documentation.pxor/por or similar combination at 2 cycles latency can be used instead.The movq however executes in the float unit, thereby saving MMX executionresources.  With the right juggling, data moves shouldn't be on a dependentchain.L1 is write-through, but the write-combining sounds like it does enough tonot require explicit destination prefetching.xmm registers so far haven't found a use, but not much effort has beenexpended.  A configure test for whether the operating system knowsfxsave/fxrestor will be needed if they're used.REFERENCESIntel Pentium-4 processor manuals,	http://developer.intel.com/design/pentium4/manuals"Intel Pentium 4 Processor Optimization Reference Manual", Intel, 2001,order number 248966.  Available on-line:	http://developer.intel.com/design/pentium4/manuals/248966.htm----------------Local variables:mode: textfill-column: 76End:

readme - 源码说明

本页面展示了「The GNU MP Bignum Library」中的 readme 源码文件，采用编程语言编写，共 114 行代码。您可以在线阅读完整代码内容，也可以返回资源详情页下载完整源码包进行本地学习和开发。

虫虫开发者社区收录了大量与GMP库相关的技术资源，包括源代码、技术文档、电路图等，是电子工程师和嵌入式开发者的专业学习平台。

⌨️ 快捷键说明

复制代码Ctrl + C

搜索代码Ctrl + F

全屏模式F11

增大字号Ctrl + =

减小字号Ctrl + -

显示快捷键?