diff --git a/de/15.9/user/index.rst b/de/15.9/user/index.rst index 871cdcfd..dfb9c953 100644 --- a/de/15.9/user/index.rst +++ b/de/15.9/user/index.rst @@ -3,7 +3,7 @@ Wie Suchanfragen in |Fess| geschrieben werden: UND, ODER und NICHT, Feld- und Labelsuche, Sortierung, Platzhalter, Bereiche, Boosting, -unscharfe Suche und Geo-Suche. +unscharfe Suche, Näherungssuche und Geo-Suche. .. toctree:: :maxdepth: 2 @@ -19,6 +19,7 @@ unscharfe Suche und Geo-Suche. search-range search-boost search-fuzzy + search-proximity search-geo search-additional role-search diff --git a/de/15.9/user/search-fuzzy.rst b/de/15.9/user/search-fuzzy.rst index 29d33efb..658032f6 100644 --- a/de/15.9/user/search-fuzzy.rst +++ b/de/15.9/user/search-fuzzy.rst @@ -41,7 +41,7 @@ Nutzungsbedingungen Beachten Sie bei der Verwendung der unscharfen Suche die folgenden Punkte: -* Die unscharfe Suche wird auf einzelne Wörter angewendet. Sie kann nicht auf Phrasen angewendet werden, die in Anführungszeichen eingeschlossen sind. Eine Zahl, die einer Phrase angehängt wird (zum Beispiel ``"Fess Search"~2``), stellt dabei keine unscharfe Suche dar, sondern eine Näherungssuche (proximity search), die den Abstand zwischen den Wörtern angibt. +* Die unscharfe Suche wird auf einzelne Wörter angewendet. Sie kann nicht auf Phrasen angewendet werden, die in Anführungszeichen eingeschlossen sind. Eine Zahl, die einer Phrase angehängt wird (zum Beispiel ``"Fess Search"~2``), stellt dabei keine unscharfe Suche dar, sondern eine :doc:`Näherungssuche ` (proximity search), die den Abstand zwischen den Wörtern angibt. * Die unscharfe Suche erfolgt auf Basis der im Index registrierten Wörter, wobei der Suchbegriff nicht erneut analysiert wird. Daher funktioniert sie möglicherweise nicht wie erwartet bei Texten wie Japanisch, die per Bi-Gramm oder morphologischer Analyse in Token zerlegt werden. Die unscharfe Suche ist vor allem bei alphanumerischen Wörtern wirksam. * Bei sehr kurzen Wörtern mit 1 bis 2 Zeichen kann das Verhalten trotz angehängtem "~" einer exakten Übereinstimmung nahekommen, da ein Treffer nur zustande kommt, wenn die Editierdistanz kleiner als die Wortlänge ist. @@ -61,4 +61,5 @@ Siehe auch ========== - :doc:`search-wildcard` +- :doc:`search-proximity` - :doc:`special-char` diff --git a/de/15.9/user/search-proximity.rst b/de/15.9/user/search-proximity.rst new file mode 100644 index 00000000..24cd16f6 --- /dev/null +++ b/de/15.9/user/search-proximity.rst @@ -0,0 +1,70 @@ +================ +Näherungssuche +================ +Näherungssuche (Proximity-Suche) +================================== + +Mit der Näherungssuche finden Sie Dokumente, in denen die Wörter einer Phrase nahe beieinander stehen, auch wenn sie nicht direkt aufeinander folgen. Sie ist nützlich, wenn zwischen den gesuchten Wörtern weitere Wörter stehen können. + +Verwendung +------------ + +Schließen Sie die Wörter in doppelte Anführungszeichen ein und fügen Sie hinter dem schließenden Anführungszeichen "~" und eine Zahl an. + +Die folgende Suche findet beispielsweise Dokumente, in denen "Fess" und "Suche" innerhalb eines Abstands von 3 vorkommen: + +:: + + "Fess Suche"~3 + +Die Zahl gibt die maximale Anzahl an Positionsverschiebungen an, die zwischen den Wörtern zulässig sind (siehe "Wie der Abstand gezählt wird" weiter unten). Je größer die Zahl, desto weiter dürfen die Wörter auseinander liegen. + +Sie können die Näherungssuche auch auf ein bestimmtes Feld anwenden. Im folgenden Beispiel wird das Feld title durchsucht. + +:: + + title:"Fess Suche"~3 + +Wenn Sie die Zahl weglassen und nur "~" angeben (zum Beispiel ``"Fess Suche"~``), wird die Phrase als normale Phrasensuche behandelt, bei der die Wörter direkt aufeinander folgen müssen. Eine Dezimalzahl wird auf eine ganze Zahl gekürzt (``~2.5`` wird als ``~2`` behandelt). + +Die Näherungssuche kann mit einem Boost kombiniert werden. Im folgenden Beispiel wird die Näherungssuche mit dem Faktor 2 gewichtet (siehe :doc:`search-boost`). + +:: + + "Fess Suche"~5^2 + +Wie der Abstand gezählt wird +------------------------------ + +* Die Zahl ist die maximale Anzahl an Positionsverschiebungen zwischen den Wörtern. Sie wird in Token gezählt, die der Analyzer des Zielfelds erzeugt, nicht in Zeichen oder durch Leerzeichen getrennten Wörtern. Bei englischem Text entspricht sie ungefähr der Anzahl der Wörter zwischen den Begriffen. Wörter, die bei der Analyse als Stoppwörter entfernt werden, werden dennoch mitgezählt. +* Die Wörter dürfen in anderer Reihenfolge vorkommen, für die umgekehrte Reihenfolge ist jedoch eine größere Zahl erforderlich. Ein Dokument mit "quick brown fox" trifft beispielsweise auf ``"quick fox"~1`` zu, auf ``"fox quick"`` aber erst ab ``~3``. + +:: + + "quick fox"~1 + "fox quick"~3 + +Japanische, chinesische und koreanische Texte (CJK) +----------------------------------------------------- + +Bei japanischen und anderen CJK-Texten hängt die Einheit, in der der Abstand gezählt wird, vom Feld ab. + +* In den allgemeinen Feldern title und content entspricht der Abstand in etwa der Anzahl der Zeichen zwischen den Wörtern. +* In den sprachspezifischen Feldern wird in morphologischen Token gezählt; entfernte Partikel werden ebenfalls mitgezählt. + +Trennen Sie die Wörter daher durch Leerzeichen und geben Sie eine großzügige Zahl an. + +:: + + "全文 検索"~5 + "大阪 おいしい"~10 + +Eine Zeichenfolge ohne Leerzeichen (zum Beispiel ``"大阪おいしい"~10``) funktioniert als Näherungssuche zwischen den Wörtern möglicherweise nicht wie erwartet, da sie in den allgemeinen Feldern als eine zusammenhängende Zeichenfolge abgeglichen wird. Trennen Sie die Wörter daher durch Leerzeichen. Der Abstand lässt sich nicht exakt in eine Zeichenanzahl umrechnen. Beginnen Sie daher mit einer großzügigen Zahl und grenzen Sie sie anhand der Ergebnisse ein. + +Siehe auch +============ + +- :doc:`search-fuzzy` +- :doc:`search-boost` +- :doc:`search-field` +- :doc:`special-char` diff --git a/de/15.9/user/special-char.rst b/de/15.9/user/special-char.rst index c9f147d0..429164f9 100644 --- a/de/15.9/user/special-char.rst +++ b/de/15.9/user/special-char.rst @@ -8,7 +8,7 @@ Die folgenden Zeichen haben in der Syntax der Suchanfrage eine besondere Bedeutu + - && || ! ( ) { } [ ] ^ " ~ * ? : \ / -Diese Zeichen werden verwendet, um Suchfunktionen wie erforderliche/ausgeschlossene Begriffe (``+`` ``-``), boolesche Operatoren (``&&`` ``||`` ``!``), Gruppierung (``( )``), Bereichssuche (``[ ]`` ``{ }``), Boost-Suche (``^``), Phrasensuche (``"``), unscharfe Suche (``~``), Platzhaltersuche (``*`` ``?``) und Feldsuche (``:``) aufzurufen. +Diese Zeichen werden verwendet, um Suchfunktionen wie erforderliche/ausgeschlossene Begriffe (``+`` ``-``), boolesche Operatoren (``&&`` ``||`` ``!``), Gruppierung (``( )``), Bereichssuche (``[ ]`` ``{ }``), Boost-Suche (``^``), Phrasensuche (``"``), unscharfe Suche und Näherungssuche (``~``), Platzhaltersuche (``*`` ``?``) und Feldsuche (``:``) aufzurufen. Wird beispielsweise ein in einer URL oder einem Dateipfad enthaltenes „/" oder „:" oder ein in Programmcode enthaltenes „+" oder „-" unmaskiert als Suchbegriff verwendet, kann dies zu einem unerwarteten Suchergebnis führen. @@ -43,8 +43,8 @@ Liste der Sonderzeichen und ihre Bedeutung - Phrasensuche (behandelt den eingeschlossenen Text als zusammenhängendes Wort; kann auch anstelle einer Maskierung verwendet werden) - :doc:`advanced-search` * - ``~`` - - Unscharfe Suche (Fuzzy-Suche) - - :doc:`search-fuzzy` + - Unscharfe Suche (nach einem Wort) / Näherungssuche (nach einer Phrase) + - :doc:`search-fuzzy` / :doc:`search-proximity` * - ``*`` ``?`` - Wildcard-Suche - :doc:`search-wildcard` @@ -84,4 +84,5 @@ Siehe auch - :doc:`search-field` - :doc:`search-wildcard` - :doc:`search-fuzzy` +- :doc:`search-proximity` - :doc:`advanced-search` diff --git a/en/15.9/user/index.rst b/en/15.9/user/index.rst index da6a643c..51afeaa1 100644 --- a/en/15.9/user/index.rst +++ b/en/15.9/user/index.rst @@ -2,8 +2,8 @@ ================== How to write search queries in |Fess|: AND, OR and NOT, field and -label search, sorting, wildcards, ranges, boosting, fuzzy search, and -geo search. +label search, sorting, wildcards, ranges, boosting, fuzzy search, proximity search, +and geo search. .. toctree:: :maxdepth: 2 @@ -19,6 +19,7 @@ geo search. search-range search-boost search-fuzzy + search-proximity search-geo search-additional role-search diff --git a/en/15.9/user/search-fuzzy.rst b/en/15.9/user/search-fuzzy.rst index a7a8de9d..df8b859e 100644 --- a/en/15.9/user/search-fuzzy.rst +++ b/en/15.9/user/search-fuzzy.rst @@ -41,7 +41,7 @@ Usage Conditions Please note the following points when using fuzzy search. -* Fuzzy search is applied on a per-word basis. It cannot be applied to phrases enclosed in quotation marks. Note that a number placed after a phrase (for example, ``"Fess Search"~2``) is not a fuzzy search but a proximity search that represents the distance between words. +* Fuzzy search is applied on a per-word basis. It cannot be applied to phrases enclosed in quotation marks. Note that a number placed after a phrase (for example, ``"Fess Search"~2``) is not a fuzzy search but a :doc:`proximity search ` that represents the distance between words. * Fuzzy search targets words that have been registered in the index, and the search term is not re-analyzed. As a result, it may not work as expected for text such as Japanese, which is tokenized using bi-grams or morphological analysis. Fuzzy search is mainly effective for alphanumeric words. * For very short words of one or two characters, since the edit distance must be smaller than the length of the word for a match to occur, adding "~" may result in behavior close to an exact match. @@ -61,4 +61,5 @@ Related Topics -------------- - :doc:`search-wildcard` - Wildcard search +- :doc:`search-proximity` - Proximity search - :doc:`special-char` - Special characters and escaping diff --git a/en/15.9/user/search-proximity.rst b/en/15.9/user/search-proximity.rst new file mode 100644 index 00000000..5a1f6c4c --- /dev/null +++ b/en/15.9/user/search-proximity.rst @@ -0,0 +1,70 @@ +================== +Proximity Search +================== +Proximity Search (Word Distance Search) +========================================= + +Proximity search finds documents in which the words of a phrase appear close to each other, even if they are not directly adjacent. It is useful when other words may appear between the words you are looking for. + +Usage +------- + +Enclose the words in double quotation marks, and add "~" and a number after the closing quotation mark. + +For example, the following search finds documents in which "Fess" and "search" appear within a distance of 3: + +:: + + "Fess search"~3 + +The number is the maximum number of position moves allowed between the words (see "How the Distance Is Counted" below). The larger the number, the farther apart the words may be. + +You can also perform a proximity search on a specific field. In the following example, the title field is searched. + +:: + + title:"Fess search"~3 + +If you omit the number and specify only "~" (for example, ``"Fess search"~``), the phrase is searched as a normal phrase in which the words must be adjacent. A decimal number is truncated to an integer (``~2.5`` is treated as ``~2``). + +Proximity search can be combined with a boost. In the following example, the proximity search is boosted by 2 (see :doc:`search-boost`). + +:: + + "Fess search"~5^2 + +How the Distance Is Counted +----------------------------- + +* The number is the maximum number of position moves allowed between the words. It is counted in tokens produced by the analyzer of the target field, not in characters or whitespace-separated words. For English text, it is roughly the number of words between the terms. Words removed as stop words during analysis are still counted. +* The words may appear in a different order, but a reversed order requires a larger number. For example, a document containing "quick brown fox" matches ``"quick fox"~1``, but it matches ``"fox quick"`` only with ``~3`` or larger. + +:: + + "quick fox"~1 + "fox quick"~3 + +Japanese and Other CJK Text +----------------------------- + +For Japanese and other CJK text, the unit in which the distance is counted depends on the field. + +* On the general title and content fields, the distance is close to the number of characters between the words. +* On the language-specific fields, the distance is counted in morphological tokens, and removed particles are also counted. + +Because of this, separate the words with spaces and specify a generous number. + +:: + + "全文 検索"~5 + "大阪 おいしい"~10 + +A string without spaces (for example, ``"大阪おいしい"~10``) may not work as a proximity search between the words as expected, because on the general fields it is matched as one continuous string. Separate the words with spaces. The distance cannot be converted exactly into a number of characters, so start with a generous number and narrow it down while checking the results. + +Related Topics +---------------- + +- :doc:`search-fuzzy` - Fuzzy search +- :doc:`search-boost` - Boost search +- :doc:`search-field` - Field-specified search +- :doc:`special-char` - Special characters and escaping diff --git a/en/15.9/user/special-char.rst b/en/15.9/user/special-char.rst index b1b1fa25..78c0db55 100644 --- a/en/15.9/user/special-char.rst +++ b/en/15.9/user/special-char.rst @@ -8,7 +8,7 @@ The following characters have a special meaning in the search query syntax, so t + - && || ! ( ) { } [ ] ^ " ~ * ? : \ / -These characters are used to invoke search features such as required/prohibited terms (``+`` ``-``), boolean operators (``&&`` ``||`` ``!``), grouping (``( )``), range search (``[ ]`` ``{ }``), boost search (``^``), phrase search (``"``), fuzzy search (``~``), wildcard search (``*`` ``?``), and field search (``:``). +These characters are used to invoke search features such as required/prohibited terms (``+`` ``-``), boolean operators (``&&`` ``||`` ``!``), grouping (``( )``), range search (``[ ]`` ``{ }``), boost search (``^``), phrase search (``"``), fuzzy search and proximity search (``~``), wildcard search (``*`` ``?``), and field search (``:``). For example, searching directly for a "/" or ":" in a URL or file path, or a "+" or "-" in program code, can produce unexpected search results. See below for how to escape these characters. @@ -43,8 +43,8 @@ List of Special Characters - Phrase search (treats the enclosed text as a single phrase; can also be used instead of escaping) - :doc:`advanced-search` * - ``~`` - - Fuzzy search - - :doc:`search-fuzzy` + - Fuzzy search (after a word) / proximity search (after a phrase) + - :doc:`search-fuzzy` / :doc:`search-proximity` * - ``*`` ``?`` - Wildcard search - :doc:`search-wildcard` @@ -85,4 +85,5 @@ Related Topics - :doc:`search-field` - Field-specified search - :doc:`search-wildcard` - Wildcard search - :doc:`search-fuzzy` - Fuzzy search +- :doc:`search-proximity` - Proximity search - :doc:`advanced-search` - Advanced search options diff --git a/es/15.9/user/index.rst b/es/15.9/user/index.rst index 9171371d..3750616c 100644 --- a/es/15.9/user/index.rst +++ b/es/15.9/user/index.rst @@ -3,7 +3,7 @@ Guía de Usuario de |Fess| Cómo escribir consultas de búsqueda en |Fess|: AND, OR y NOT, búsqueda por campo y por etiqueta, ordenación, comodines, rangos, boosting, -búsqueda difusa y búsqueda geográfica. +búsqueda difusa, búsqueda de proximidad y búsqueda geográfica. .. toctree:: :maxdepth: 2 @@ -19,6 +19,7 @@ búsqueda difusa y búsqueda geográfica. search-range search-boost search-fuzzy + search-proximity search-geo search-additional role-search diff --git a/es/15.9/user/search-fuzzy.rst b/es/15.9/user/search-fuzzy.rst index 857dcbef..2459b315 100644 --- a/es/15.9/user/search-fuzzy.rst +++ b/es/15.9/user/search-fuzzy.rst @@ -41,7 +41,7 @@ Condiciones de uso Tenga en cuenta los siguientes puntos al utilizar la búsqueda difusa. -* La búsqueda difusa se aplica a nivel de palabra. No se puede aplicar a frases entre comillas. Además, un número agregado después de una frase (por ejemplo, ``"Fess Search"~2``) no corresponde a una búsqueda difusa, sino a una búsqueda de proximidad que indica la distancia entre palabras. +* La búsqueda difusa se aplica a nivel de palabra. No se puede aplicar a frases entre comillas. Además, un número agregado después de una frase (por ejemplo, ``"Fess Search"~2``) no corresponde a una búsqueda difusa, sino a una :doc:`búsqueda de proximidad ` que indica la distancia entre palabras. * La búsqueda difusa se realiza sobre las palabras registradas en el índice y el término de búsqueda no se vuelve a analizar. Por ello, es posible que no funcione como se espera en textos como el japonés, que se tokeniza mediante bi-gramas o análisis morfológico. La búsqueda difusa es eficaz principalmente con palabras alfanuméricas. * En el caso de palabras muy cortas, de 1 o 2 caracteres, la distancia de edición debe ser menor que la longitud de la palabra para que se produzca una coincidencia, por lo que, aunque se agregue "~", el comportamiento puede acercarse al de una coincidencia exacta. @@ -62,4 +62,5 @@ Véase también ============= - :doc:`search-wildcard` - Búsqueda con comodines +- :doc:`search-proximity` - Búsqueda de proximidad - :doc:`special-char` - Caracteres especiales diff --git a/es/15.9/user/search-proximity.rst b/es/15.9/user/search-proximity.rst new file mode 100644 index 00000000..1d339d19 --- /dev/null +++ b/es/15.9/user/search-proximity.rst @@ -0,0 +1,70 @@ +======================== +Búsqueda de proximidad +======================== +Búsqueda de proximidad (distancia entre palabras) +=================================================== + +La búsqueda de proximidad permite encontrar documentos en los que las palabras de una frase aparecen cerca unas de otras, aunque no estén juntas. Es útil cuando pueden aparecer otras palabras entre las que se están buscando. + +Cómo utilizar +--------------- + +Encierre las palabras entre comillas dobles y agregue "~" y un número después de la comilla de cierre. + +Por ejemplo, la siguiente búsqueda encuentra documentos en los que "Fess" y "búsqueda" aparecen a una distancia de 3 o menos: + +:: + + "Fess búsqueda"~3 + +El número es la cantidad máxima de desplazamientos de posición permitidos entre las palabras (consulte "Cómo se cuenta la distancia" más adelante). Cuanto mayor sea el número, más separadas pueden estar las palabras. + +También puede realizar una búsqueda de proximidad en un campo específico. En el siguiente ejemplo, se busca en el campo title. + +:: + + title:"Fess búsqueda"~3 + +Si omite el número y especifica únicamente "~" (por ejemplo, ``"Fess búsqueda"~``), la frase se busca como una frase normal en la que las palabras deben ser contiguas. Un número decimal se trunca a un entero (``~2.5`` se trata como ``~2``). + +La búsqueda de proximidad se puede combinar con un impulso (boost). En el siguiente ejemplo, la búsqueda de proximidad se multiplica por 2 (consulte :doc:`search-boost`). + +:: + + "Fess búsqueda"~5^2 + +Cómo se cuenta la distancia +----------------------------- + +* El número es la cantidad máxima de desplazamientos de posición permitidos entre las palabras. Se cuenta en tokens generados por el analizador del campo de destino, no en caracteres ni en palabras separadas por espacios. En textos en inglés, equivale aproximadamente al número de palabras que hay entre los términos. Las palabras eliminadas como palabras vacías (stop words) durante el análisis también se cuentan. +* Las palabras pueden aparecer en otro orden, pero el orden inverso requiere un número mayor. Por ejemplo, un documento que contiene "quick brown fox" coincide con ``"quick fox"~1``, pero solo coincide con ``"fox quick"`` a partir de ``~3``. + +:: + + "quick fox"~1 + "fox quick"~3 + +Texto en japonés, chino y coreano (CJK) +----------------------------------------- + +En textos en japonés y otros textos CJK, la unidad en que se cuenta la distancia depende del campo. + +* En los campos generales title y content, la distancia es cercana al número de caracteres que hay entre las palabras. +* En los campos específicos de cada idioma, la distancia se cuenta en tokens morfológicos, y las partículas eliminadas también se cuentan. + +Por ello, separe las palabras con espacios y especifique un número generoso. + +:: + + "全文 検索"~5 + "大阪 おいしい"~10 + +Es posible que una cadena sin espacios (por ejemplo, ``"大阪おいしい"~10``) no funcione como se espera como búsqueda de proximidad entre palabras, ya que en los campos generales se compara como una única cadena continua. Separe las palabras con espacios. La distancia no se puede convertir exactamente en un número de caracteres, por lo que conviene empezar con un número generoso y reducirlo a medida que comprueba los resultados. + +Véase también +=============== + +- :doc:`search-fuzzy` - Búsqueda difusa +- :doc:`search-boost` - Búsqueda con boost +- :doc:`search-field` - Búsqueda con especificación de campos +- :doc:`special-char` - Caracteres especiales diff --git a/es/15.9/user/special-char.rst b/es/15.9/user/special-char.rst index 9254d65e..370e70d0 100644 --- a/es/15.9/user/special-char.rst +++ b/es/15.9/user/special-char.rst @@ -8,7 +8,7 @@ Los siguientes caracteres tienen un significado especial en la sintaxis de las c + - && || ! ( ) { } [ ] ^ " ~ * ? : \ / -Estos caracteres se utilizan para invocar funciones de búsqueda como términos obligatorios/prohibidos (``+`` ``-``), operadores booleanos (``&&`` ``||`` ``!``), agrupación (``( )``), búsqueda por rango (``[ ]`` ``{ }``), búsqueda con impulso (boost) (``^``), búsqueda de frases (``"``), búsqueda difusa (fuzzy) (``~``), búsqueda con comodines (``*`` ``?``) y búsqueda por campo (``:``). +Estos caracteres se utilizan para invocar funciones de búsqueda como términos obligatorios/prohibidos (``+`` ``-``), operadores booleanos (``&&`` ``||`` ``!``), agrupación (``( )``), búsqueda por rango (``[ ]`` ``{ }``), búsqueda con impulso (boost) (``^``), búsqueda de frases (``"``), búsqueda difusa (fuzzy) y búsqueda de proximidad (``~``), búsqueda con comodines (``*`` ``?``) y búsqueda por campo (``:``). Por ejemplo, si busca directamente símbolos como "/" o ":" incluidos en una URL o ruta de archivo, o "+" o "-" que aparecen en código de programación, puede obtener resultados de búsqueda no deseados. Consulte a continuación el método de escape. @@ -44,8 +44,8 @@ Lista de caracteres especiales y su significado - Búsqueda de frases (trata el texto entre comillas como una sola unidad; también puede usarse en lugar del escape) - :doc:`advanced-search` * - ``~`` - - Búsqueda difusa (búsqueda aproximada) - - :doc:`search-fuzzy` + - Búsqueda difusa (tras una palabra) / búsqueda de proximidad (tras una frase) + - :doc:`search-fuzzy` / :doc:`search-proximity` * - ``*`` ``?`` - Búsqueda con comodines - :doc:`search-wildcard` @@ -87,4 +87,5 @@ Véase también - :doc:`search-field` - Búsqueda con especificación de campos - :doc:`search-wildcard` - Búsqueda con comodines - :doc:`search-fuzzy` - Búsqueda difusa +- :doc:`search-proximity` - Búsqueda de proximidad - :doc:`advanced-search` - Búsqueda avanzada diff --git a/fr/15.9/user/index.rst b/fr/15.9/user/index.rst index 85c6e527..fe19ad79 100644 --- a/fr/15.9/user/index.rst +++ b/fr/15.9/user/index.rst @@ -3,7 +3,7 @@ Guide utilisateur |Fess| Comment écrire des requêtes dans |Fess| : AND, OR et NOT, recherche par champ et par étiquette, tri, jokers, plages, boost, recherche -floue et recherche géographique. +floue, recherche de proximité et recherche géographique. .. toctree:: :maxdepth: 2 @@ -19,6 +19,7 @@ floue et recherche géographique. search-range search-boost search-fuzzy + search-proximity search-geo search-additional role-search diff --git a/fr/15.9/user/search-fuzzy.rst b/fr/15.9/user/search-fuzzy.rst index 4703a4c1..bf5be029 100644 --- a/fr/15.9/user/search-fuzzy.rst +++ b/fr/15.9/user/search-fuzzy.rst @@ -41,7 +41,7 @@ Conditions d'utilisation Lors de l'utilisation de la recherche floue, tenez compte des points suivants : -* La recherche floue s'applique au niveau du mot. Elle ne peut pas être appliquée à une phrase entourée de guillemets. Notez que le chiffre placé après une phrase (par exemple ``"Fess Search"~2``) ne correspond pas à une recherche floue, mais à une recherche de proximité indiquant la distance entre les mots. +* La recherche floue s'applique au niveau du mot. Elle ne peut pas être appliquée à une phrase entourée de guillemets. Notez que le chiffre placé après une phrase (par exemple ``"Fess Search"~2``) ne correspond pas à une recherche floue, mais à une :doc:`recherche de proximité ` indiquant la distance entre les mots. * La recherche floue porte sur les mots enregistrés dans l'index, et le terme de recherche n'est pas réanalysé. Par conséquent, elle peut ne pas fonctionner comme prévu pour des textes tels que le japonais, qui sont tokenisés par bi-gramme ou par analyse morphologique. La recherche floue est principalement efficace pour les mots alphanumériques. * Pour les mots très courts, comme ceux de 1 à 2 caractères, la correspondance n'est possible que si la distance d'édition est inférieure à la longueur du mot ; l'ajout de « ~ » peut donc se comporter presque comme une correspondance exacte. @@ -61,4 +61,5 @@ Voir aussi ========== - :doc:`search-wildcard` +- :doc:`search-proximity` - :doc:`special-char` diff --git a/fr/15.9/user/search-proximity.rst b/fr/15.9/user/search-proximity.rst new file mode 100644 index 00000000..c59cf7cb --- /dev/null +++ b/fr/15.9/user/search-proximity.rst @@ -0,0 +1,70 @@ +======================== +Recherche de proximité +======================== +Recherche de proximité (distance entre les mots) +================================================== + +La recherche de proximité permet de trouver les documents dans lesquels les mots d'une phrase apparaissent proches les uns des autres, même s'ils ne sont pas directement adjacents. Elle est utile lorsque d'autres mots peuvent s'intercaler entre les mots recherchés. + +Utilisation +------------- + +Entourez les mots de guillemets doubles, puis ajoutez « ~ » et un chiffre après le guillemet fermant. + +Par exemple, la recherche suivante trouve les documents dans lesquels « Fess » et « recherche » apparaissent à une distance de 3 au plus : + +:: + + "Fess recherche"~3 + +Le nombre correspond au nombre maximal de déplacements de position autorisés entre les mots (voir « Comment la distance est comptée » ci-dessous). Plus le nombre est grand, plus les mots peuvent être éloignés. + +Vous pouvez également effectuer une recherche de proximité sur un champ précis. Dans l'exemple suivant, la recherche porte sur le champ title. + +:: + + title:"Fess recherche"~3 + +Si vous omettez le nombre et n'indiquez que « ~ » (par exemple ``"Fess recherche"~``), la phrase est recherchée comme une phrase normale dont les mots doivent être adjacents. Un nombre décimal est tronqué à sa partie entière (``~2.5`` est traité comme ``~2``). + +La recherche de proximité peut être combinée avec un boost. Dans l'exemple suivant, la recherche de proximité est pondérée par 2 (voir :doc:`search-boost`). + +:: + + "Fess recherche"~5^2 + +Comment la distance est comptée +--------------------------------- + +* Le nombre correspond au nombre maximal de déplacements de position autorisés entre les mots. Il est compté en tokens produits par l'analyseur du champ cible, et non en caractères ni en mots séparés par des espaces. Pour un texte en anglais, il correspond approximativement au nombre de mots situés entre les termes. Les mots supprimés comme mots vides (stop words) lors de l'analyse sont également comptés. +* Les mots peuvent apparaître dans un ordre différent, mais l'ordre inversé nécessite un nombre plus grand. Par exemple, un document contenant « quick brown fox » correspond à ``"quick fox"~1``, mais ne correspond à ``"fox quick"`` qu'à partir de ``~3``. + +:: + + "quick fox"~1 + "fox quick"~3 + +Texte japonais, chinois et coréen (CJK) +----------------------------------------- + +Pour le japonais et les autres textes CJK, l'unité dans laquelle la distance est comptée dépend du champ. + +* Dans les champs généraux title et content, la distance est proche du nombre de caractères situés entre les mots. +* Dans les champs propres à chaque langue, la distance est comptée en tokens morphologiques, et les particules supprimées sont également comptées. + +Séparez donc les mots par des espaces et indiquez un nombre généreux. + +:: + + "全文 検索"~5 + "大阪 おいしい"~10 + +Une chaîne sans espaces (par exemple ``"大阪おいしい"~10``) peut ne pas fonctionner comme prévu en tant que recherche de proximité entre les mots, car dans les champs généraux elle est comparée comme une seule chaîne continue. Séparez les mots par des espaces. La distance ne peut pas être convertie exactement en nombre de caractères ; commencez donc par un nombre généreux, puis affinez-le en vérifiant les résultats. + +Voir aussi +============ + +- :doc:`search-fuzzy` +- :doc:`search-boost` +- :doc:`search-field` +- :doc:`special-char` diff --git a/fr/15.9/user/special-char.rst b/fr/15.9/user/special-char.rst index b32e9325..edfc077a 100644 --- a/fr/15.9/user/special-char.rst +++ b/fr/15.9/user/special-char.rst @@ -8,7 +8,7 @@ Les caractères suivants ont une signification particulière dans la syntaxe des + - && || ! ( ) { } [ ] ^ " ~ * ? : \ / -Ces caractères permettent d'invoquer des fonctionnalités de recherche telles que les termes obligatoires/interdits (``+`` ``-``), les opérateurs booléens (``&&`` ``||`` ``!``), le regroupement (``( )``), la recherche par plage (``[ ]`` ``{ }``), la recherche avec pondération (``^``), la recherche de phrase (``"``), la recherche floue (``~``), la recherche par caractère générique (``*`` ``?``) et la recherche par champ (``:``). +Ces caractères permettent d'invoquer des fonctionnalités de recherche telles que les termes obligatoires/interdits (``+`` ``-``), les opérateurs booléens (``&&`` ``||`` ``!``), le regroupement (``( )``), la recherche par plage (``[ ]`` ``{ }``), la recherche avec pondération (``^``), la recherche de phrase (``"``), la recherche floue et la recherche de proximité (``~``), la recherche par caractère générique (``*`` ``?``) et la recherche par champ (``:``). Par exemple, si vous recherchez tels quels des caractères comme ``/`` ou ``:`` présents dans une URL ou un chemin de fichier, ou encore ``+`` ou ``-`` utilisés dans du code, vous risquez d'obtenir des résultats de recherche inattendus. Reportez-vous à la section ci-dessous pour la méthode d'échappement à utiliser. @@ -43,8 +43,8 @@ Liste des caractères spéciaux et leur signification - Recherche de phrase (le texte entouré de guillemets est traité comme une seule expression ; peut aussi servir à échapper des caractères spéciaux) - :doc:`advanced-search` * - ``~`` - - Recherche floue (recherche approximative) - - :doc:`search-fuzzy` + - Recherche floue (après un mot) / recherche de proximité (après une phrase) + - :doc:`search-fuzzy` / :doc:`search-proximity` * - ``*`` ``?`` - Recherche avec caractères génériques - :doc:`search-wildcard` @@ -84,4 +84,5 @@ Voir aussi - :doc:`search-field` - :doc:`search-wildcard` - :doc:`search-fuzzy` +- :doc:`search-proximity` - :doc:`advanced-search` diff --git a/ja/15.9/user/index.rst b/ja/15.9/user/index.rst index e7085a89..2e4dcd64 100644 --- a/ja/15.9/user/index.rst +++ b/ja/15.9/user/index.rst @@ -2,7 +2,7 @@ ================== AND・OR・NOT、フィールド検索やラベル検索、並び替え、ワイルドカード、\ -範囲指定、ブースト、あいまい検索、位置情報検索など、|Fess| での\ +範囲指定、ブースト、あいまい検索、近接検索、位置情報検索など、|Fess| での\ 検索の書き方を説明します。 .. toctree:: @@ -19,6 +19,7 @@ AND・OR・NOT、フィールド検索やラベル検索、並び替え、ワイ search-range search-boost search-fuzzy + search-proximity search-geo search-additional role-search diff --git a/ja/15.9/user/search-fuzzy.rst b/ja/15.9/user/search-fuzzy.rst index 9266b6b5..eaa25359 100644 --- a/ja/15.9/user/search-fuzzy.rst +++ b/ja/15.9/user/search-fuzzy.rst @@ -41,7 +41,7 @@ 曖昧検索を利用する際は以下の点に注意してください。 -* 曖昧検索は単語単位で適用されます。引用符で囲んだフレーズには適用できません。なお、フレーズの後ろに付けた数字 (たとえば ``"Fess Search"~2``) は、曖昧検索ではなく単語間の距離を表す近接検索になります。 +* 曖昧検索は単語単位で適用されます。引用符で囲んだフレーズには適用できません。なお、フレーズの後ろに付けた数字 (たとえば ``"Fess Search"~2``) は、曖昧検索ではなく単語間の距離を表す\ :doc:`近接検索 `\ になります。 * 曖昧検索はインデックスに登録された単語を対象に行われ、検索語は再解析されません。そのため、bi-gram や形態素解析でトークン化される日本語などのテキストでは期待どおりに動作しないことがあります。曖昧検索は主に英数字の単語で有効です。 * 1〜2 文字などの非常に短い単語は、編集距離が単語の長さより小さくないと一致しないため、「~」を付けても完全一致に近い動作になる場合があります。 @@ -62,4 +62,5 @@ ======== - :doc:`search-wildcard` +- :doc:`search-proximity` - :doc:`special-char` diff --git a/ja/15.9/user/search-proximity.rst b/ja/15.9/user/search-proximity.rst new file mode 100644 index 00000000..ec3ec634 --- /dev/null +++ b/ja/15.9/user/search-proximity.rst @@ -0,0 +1,70 @@ +========== +近接検索 +========== +近接検索(プロキシミティ検索) +============================== + +近接検索は、フレーズに含まれる単語が隣り合っていなくても、互いに近い位置にあるドキュメントを検索する方法です。探している単語の間に別の語が入る可能性がある場合に有効です。 + +利用方法 +---------- + +単語を二重引用符で囲み、閉じる引用符の後ろに「~」と数字を付加します。 + +たとえば、以下のように入力すると、「Fess」と「検索」が距離 3 以内にあるドキュメントを検索できます。 + +:: + + "Fess 検索"~3 + +数字は、単語の間で許容される位置の移動回数の最大値です (詳しくは後述の「距離の数え方」を参照してください)。数字が大きいほど、単語が離れていても一致します。 + +フィールドを指定して近接検索を行うこともできます。次の例では、title フィールドを検索します。 + +:: + + title:"Fess 検索"~3 + +数字を省略して「~」のみを指定した場合 (たとえば ``"Fess 検索"~``) は、単語が隣り合っている必要がある通常のフレーズ検索として扱われます。小数を指定した場合は小数点以下が切り捨てられます (``~2.5`` は ``~2`` として扱われます)。 + +ブーストと組み合わせることもできます。次の例では、近接検索に 2 のブーストを適用します (:doc:`search-boost` を参照)。 + +:: + + "Fess 検索"~5^2 + +距離の数え方 +-------------- + +* 数字は、単語の間で許容される位置の移動回数の最大値です。対象フィールドのアナライザーが生成したトークンを単位として数え、文字数や空白で区切られた単語数ではありません。英語のテキストでは、おおむね語の間に挟まる単語数になります。解析時にストップワードとして除去された語も数えられます。 +* 単語の順序が入れ替わっていても一致しますが、逆順の場合はより大きな数字が必要です。たとえば「quick brown fox」を含むドキュメントは ``"quick fox"~1`` に一致しますが、``"fox quick"`` に一致させるには ``~3`` 以上が必要です。 + +:: + + "quick fox"~1 + "fox quick"~3 + +日本語などの CJK テキストの場合 +--------------------------------- + +日本語などの CJK テキストでは、距離を数える単位がフィールドによって異なります。 + +* title や content などの汎用フィールドでは、単語の間にある文字数に近い値になります。 +* 言語別のフィールドでは、形態素を単位として数えられ、除去された助詞なども数えられます。 + +そのため、単語は空白で区切り、数字は余裕を持って指定してください。 + +:: + + "全文 検索"~5 + "大阪 おいしい"~10 + +空白を含まない文字列 (たとえば ``"大阪おいしい"~10``) は、汎用フィールドでは連続した 1 つの文字列として照合されるため、単語間の近接検索として期待どおりに動作しないことがあります。単語は空白で区切ってください。距離を文字数に正確に換算することはできないため、まず大きめの数字を指定し、結果を確認しながら絞り込んでください。 + +関連項目 +========== + +- :doc:`search-fuzzy` +- :doc:`search-boost` +- :doc:`search-field` +- :doc:`special-char` diff --git a/ja/15.9/user/special-char.rst b/ja/15.9/user/special-char.rst index ab0c7de7..a864348f 100644 --- a/ja/15.9/user/special-char.rst +++ b/ja/15.9/user/special-char.rst @@ -8,7 +8,7 @@ + - && || ! ( ) { } [ ] ^ " ~ * ? : \ / -これらの文字は、必須・除外指定(``+`` ``-``)、ブール演算(``&&`` ``||`` ``!``)、グループ化(``( )``)、範囲検索(``[ ]`` ``{ }``)、ブースト検索(``^``)、フレーズ検索(``"``)、あいまい検索(``~``)、ワイルドカード検索(``*`` ``?``)、フィールド指定検索(``:``)などの検索機能を呼び出すために使われます。 +これらの文字は、必須・除外指定(``+`` ``-``)、ブール演算(``&&`` ``||`` ``!``)、グループ化(``( )``)、範囲検索(``[ ]`` ``{ }``)、ブースト検索(``^``)、フレーズ検索(``"``)、あいまい検索・近接検索(``~``)、ワイルドカード検索(``*`` ``?``)、フィールド指定検索(``:``)などの検索機能を呼び出すために使われます。 たとえば URL やファイルパスに含まれる「/」「:」、プログラム中の「+」「-」などをそのまま検索すると、意図しない検索結果になることがあります。エスケープの方法は以下を参照してください。 @@ -44,8 +44,8 @@ - フレーズ検索(囲んだ範囲を1語句として扱う。エスケープの代わりにも使える) - :doc:`advanced-search` * - ``~`` - - あいまい検索(ファジー検索) - - :doc:`search-fuzzy` + - あいまい検索(単語の後ろ)/ 近接検索(フレーズの後ろ) + - :doc:`search-fuzzy` / :doc:`search-proximity` * - ``*`` ``?`` - ワイルドカード検索 - :doc:`search-wildcard` @@ -87,4 +87,5 @@ - :doc:`search-field` - :doc:`search-wildcard` - :doc:`search-fuzzy` +- :doc:`search-proximity` - :doc:`advanced-search` diff --git a/ko/15.9/user/index.rst b/ko/15.9/user/index.rst index 3232e7a0..1aa4c62d 100644 --- a/ko/15.9/user/index.rst +++ b/ko/15.9/user/index.rst @@ -2,7 +2,7 @@ ================== AND·OR·NOT, 필드 검색과 라벨 검색, 정렬, 와일드카드, 범위 지정, -부스트, 유사 검색, 위치 검색 등 |Fess|\ 에서 검색어를 작성하는 방법을 +부스트, 유사 검색, 근접 검색, 위치 검색 등 |Fess|\ 에서 검색어를 작성하는 방법을 설명합니다. .. toctree:: @@ -19,6 +19,7 @@ AND·OR·NOT, 필드 검색과 라벨 검색, 정렬, 와일드카드, 범위 search-range search-boost search-fuzzy + search-proximity search-geo search-additional role-search diff --git a/ko/15.9/user/search-fuzzy.rst b/ko/15.9/user/search-fuzzy.rst index dc5d4a47..ae2f6ddb 100644 --- a/ko/15.9/user/search-fuzzy.rst +++ b/ko/15.9/user/search-fuzzy.rst @@ -41,7 +41,7 @@ 유사 검색을 사용할 때는 다음 사항에 주의하십시오. -* 유사 검색은 단어 단위로 적용됩니다. 따옴표로 묶은 구문에는 적용할 수 없습니다. 또한 구문 뒤에 붙인 숫자 (예: ``"Fess Search"~2``)는 유사 검색이 아니라 단어 사이의 거리를 나타내는 근접 검색이 됩니다. +* 유사 검색은 단어 단위로 적용됩니다. 따옴표로 묶은 구문에는 적용할 수 없습니다. 또한 구문 뒤에 붙인 숫자 (예: ``"Fess Search"~2``)는 유사 검색이 아니라 단어 사이의 거리를 나타내는 :doc:`근접 검색 `\ 이 됩니다. * 유사 검색은 인덱스에 등록된 단어를 대상으로 수행되며, 검색어는 다시 분석되지 않습니다. 따라서 bi-gram이나 형태소 분석으로 토큰화되는 일본어 등의 텍스트에서는 예상대로 동작하지 않을 수 있습니다. 유사 검색은 주로 영숫자 단어에 유효합니다. * 1~2 문자 등 매우 짧은 단어는 편집 거리가 단어 길이보다 작아야만 일치하므로, "~"를 붙여도 완전 일치에 가까운 동작이 되는 경우가 있습니다. @@ -62,4 +62,5 @@ ================= - :doc:`search-wildcard` - 와일드카드 검색 +- :doc:`search-proximity` - 근접 검색 - :doc:`special-char` - 특수 문자 diff --git a/ko/15.9/user/search-proximity.rst b/ko/15.9/user/search-proximity.rst new file mode 100644 index 00000000..e0569707 --- /dev/null +++ b/ko/15.9/user/search-proximity.rst @@ -0,0 +1,70 @@ +=========== +근접 검색 +=========== +근접 검색(단어 간 거리 검색) +============================== + +근접 검색은 구문에 포함된 단어가 서로 붙어 있지 않더라도 가까운 위치에 있는 문서를 검색하는 방법입니다. 찾으려는 단어 사이에 다른 단어가 들어갈 수 있는 경우에 유용합니다. + +사용 방법 +----------- + +단어를 큰따옴표로 묶고, 닫는 따옴표 뒤에 "~"와 숫자를 추가합니다. + +예를 들어 다음과 같이 입력하면 "Fess"와 "검색"이 거리 3 이내에 있는 문서를 검색할 수 있습니다. + +:: + + "Fess 검색"~3 + +숫자는 단어 사이에서 허용되는 위치 이동 횟수의 최댓값입니다 (자세한 내용은 아래의 "거리 계산 방식"을 참조하십시오). 숫자가 클수록 단어가 더 멀리 떨어져 있어도 일치합니다. + +필드를 지정하여 근접 검색을 수행할 수도 있습니다. 다음 예에서는 title 필드를 검색합니다. + +:: + + title:"Fess 검색"~3 + +숫자를 생략하고 "~"만 지정한 경우 (예: ``"Fess 검색"~``)에는 단어가 인접해 있어야 하는 일반 구문 검색으로 처리됩니다. 소수를 지정한 경우 소수점 이하는 버려집니다 (``~2.5``\ 는 ``~2``\ 로 처리됩니다). + +부스트와 함께 사용할 수도 있습니다. 다음 예에서는 근접 검색에 2의 부스트를 적용합니다 (:doc:`search-boost` 참조). + +:: + + "Fess 검색"~5^2 + +거리 계산 방식 +---------------- + +* 숫자는 단어 사이에서 허용되는 위치 이동 횟수의 최댓값입니다. 대상 필드의 애널라이저가 생성한 토큰을 단위로 계산하며, 문자 수나 공백으로 구분된 단어 수가 아닙니다. 영어 텍스트에서는 대략 단어 사이에 끼어 있는 단어 수에 해당합니다. 분석 시 불용어로 제거된 단어도 계산에 포함됩니다. +* 단어의 순서가 바뀌어도 일치하지만, 역순인 경우에는 더 큰 숫자가 필요합니다. 예를 들어 "quick brown fox"를 포함하는 문서는 ``"quick fox"~1``\ 에 일치하지만, ``"fox quick"``\ 에 일치시키려면 ``~3`` 이상이 필요합니다. + +:: + + "quick fox"~1 + "fox quick"~3 + +한국어 등 CJK 텍스트의 경우 +----------------------------- + +일본어, 중국어, 한국어 등 CJK 텍스트에서는 거리를 계산하는 단위가 필드에 따라 다릅니다. + +* title, content 등의 범용 필드에서는 단어 사이에 있는 문자 수에 가까운 값이 됩니다. +* 언어별 필드에서는 형태소를 단위로 계산되며, 제거된 조사 등도 계산에 포함됩니다. + +따라서 단어는 공백으로 구분하고, 숫자는 여유 있게 지정하십시오. + +:: + + "전문 검색"~5 + "오사카 맛있는"~10 + +공백이 없는 문자열 (예: ``"오사카맛있는"~10``)은 범용 필드에서는 연속된 하나의 문자열로 대조되므로, 단어 간 근접 검색으로 기대한 대로 동작하지 않을 수 있습니다. 단어는 공백으로 구분하십시오. 거리를 문자 수로 정확하게 환산할 수는 없으므로, 먼저 큰 숫자를 지정하고 결과를 확인하면서 좁혀 나가십시오. + +관련 항목 +=========== + +- :doc:`search-fuzzy` - 유사 검색(퍼지 검색) +- :doc:`search-boost` - 부스트 검색 +- :doc:`search-field` - 필드 지정 검색 +- :doc:`special-char` - 특수 문자 diff --git a/ko/15.9/user/special-char.rst b/ko/15.9/user/special-char.rst index 3a06e375..75210561 100644 --- a/ko/15.9/user/special-char.rst +++ b/ko/15.9/user/special-char.rst @@ -8,7 +8,7 @@ + - && || ! ( ) { } [ ] ^ " ~ * ? : \ / -이러한 문자는 필수/제외 지정(``+`` ``-``), 불리언 연산(``&&`` ``||`` ``!``), 그룹화(``( )``), 범위 검색(``[ ]`` ``{ }``), 부스트 검색(``^``), 프레이즈 검색(``"``), 퍼지 검색(``~``), 와일드카드 검색(``*`` ``?``), 필드 지정 검색(``:``) 등의 검색 기능을 호출하는 데 사용됩니다. +이러한 문자는 필수/제외 지정(``+`` ``-``), 불리언 연산(``&&`` ``||`` ``!``), 그룹화(``( )``), 범위 검색(``[ ]`` ``{ }``), 부스트 검색(``^``), 프레이즈 검색(``"``), 퍼지 검색·근접 검색(``~``), 와일드카드 검색(``*`` ``?``), 필드 지정 검색(``:``) 등의 검색 기능을 호출하는 데 사용됩니다. 예를 들어 URL이나 파일 경로에 포함된 "/", ":", 프로그램 코드에 포함된 "+", "-" 등을 그대로 검색하면 의도하지 않은 검색 결과가 나올 수 있습니다. 이스케이프 방법은 아래를 참조하십시오. @@ -44,8 +44,8 @@ - 프레이즈 검색(감싼 범위를 하나의 어구로 취급. 이스케이프 대신 사용 가능) - :doc:`advanced-search` * - ``~`` - - 유사 검색(퍼지 검색) - - :doc:`search-fuzzy` + - 유사 검색(단어 뒤) / 근접 검색(구문 뒤) + - :doc:`search-fuzzy` / :doc:`search-proximity` * - ``*`` ``?`` - 와일드카드 검색 - :doc:`search-wildcard` @@ -87,4 +87,5 @@ - :doc:`search-field` - 필드 지정 검색 - :doc:`search-wildcard` - 와일드카드 검색 - :doc:`search-fuzzy` - 유사 검색(퍼지 검색) +- :doc:`search-proximity` - 근접 검색 - :doc:`advanced-search` - 고급 검색 diff --git a/zh-cn/15.9/user/index.rst b/zh-cn/15.9/user/index.rst index 4f0218cd..eb6debc5 100644 --- a/zh-cn/15.9/user/index.rst +++ b/zh-cn/15.9/user/index.rst @@ -2,7 +2,7 @@ ================== 说明在 |Fess| 中编写搜索条件的方法,涵盖 AND、OR、NOT、\ -字段搜索与标签搜索、排序、通配符、范围、加权、模糊搜索和地理搜索。 +字段搜索与标签搜索、排序、通配符、范围、加权、模糊搜索、邻近搜索和地理搜索。 .. toctree:: :maxdepth: 2 @@ -18,6 +18,7 @@ search-range search-boost search-fuzzy + search-proximity search-geo search-additional role-search diff --git a/zh-cn/15.9/user/search-fuzzy.rst b/zh-cn/15.9/user/search-fuzzy.rst index c4bf63e5..7c153067 100644 --- a/zh-cn/15.9/user/search-fuzzy.rst +++ b/zh-cn/15.9/user/search-fuzzy.rst @@ -41,7 +41,7 @@ 使用模糊检索时,请注意以下几点。 -* 模糊检索以单词为单位应用,无法应用于用引号括起来的短语。此外,短语后面添加的数字(例如 ``"Fess Search"~2``)不是模糊检索,而是表示单词间距离的邻近检索。 +* 模糊检索以单词为单位应用,无法应用于用引号括起来的短语。此外,短语后面添加的数字(例如 ``"Fess Search"~2``)不是模糊检索,而是表示单词间距离的\ :doc:`邻近检索 `\ 。 * 模糊检索针对索引中已登记的单词进行,检索词不会被重新解析。因此,对于通过 bi-gram 或形态素解析进行分词的日语等文本,可能无法按预期工作。模糊检索主要对英数字单词有效。 * 对于 1~2 个字符等非常短的单词,由于编辑距离必须小于单词长度才能匹配,因此即使添加"~",也可能出现接近完全匹配的行为。 @@ -61,4 +61,5 @@ ------ - :doc:`search-wildcard` - 通配符检索 +- :doc:`search-proximity` - 邻近检索 - :doc:`special-char` - 特殊字符与转义 diff --git a/zh-cn/15.9/user/search-proximity.rst b/zh-cn/15.9/user/search-proximity.rst new file mode 100644 index 00000000..a09a2c2f --- /dev/null +++ b/zh-cn/15.9/user/search-proximity.rst @@ -0,0 +1,70 @@ +========== +邻近检索 +========== +邻近检索(单词距离检索) +========================== + +邻近检索是指,即使短语中的单词没有紧挨在一起,只要它们的位置相互靠近,也能检索到相应文档的方法。当要查找的单词之间可能夹有其他词语时,该方法非常有用。 + +使用方法 +---------- + +用双引号将单词括起来,并在右侧引号后面添加"~"和数字。 + +例如,按如下方式输入,可以检索"Fess"与"检索"之间距离在 3 以内的文档。 + +:: + + "Fess 检索"~3 + +数字表示单词之间允许的位置移动次数的最大值(详见下文"距离的计算方式")。数字越大,允许单词相隔得越远。 + +也可以指定字段进行邻近检索。下面的示例将检索 title 字段。 + +:: + + title:"Fess 检索"~3 + +如果省略数字,仅指定"~"(例如 ``"Fess 检索"~``),则按要求单词相邻的普通短语检索处理。如果指定小数,小数部分将被舍去(``~2.5`` 按 ``~2`` 处理)。 + +也可以与权重(boost)组合使用。下面的示例对邻近检索应用 2 倍权重(请参见 :doc:`search-boost`)。 + +:: + + "Fess 检索"~5^2 + +距离的计算方式 +---------------- + +* 数字表示单词之间允许的位置移动次数的最大值。它以目标字段的分析器生成的词元(token)为单位计数,而不是字符数或以空格分隔的单词数。对于英文文本,大致相当于词语之间夹着的单词数。分析时作为停用词被去除的单词也会被计入。 +* 单词的顺序可以不同,但顺序颠倒时需要更大的数字。例如,包含"quick brown fox"的文档可以匹配 ``"quick fox"~1``,但要匹配 ``"fox quick"``,则需要 ``~3`` 或更大的数字。 + +:: + + "quick fox"~1 + "fox quick"~3 + +中文等 CJK 文本 +----------------- + +对于中文、日语、韩语等 CJK 文本,计算距离的单位因字段而异。 + +* 在 title、content 等通用字段中,距离接近于单词之间的字符数。 +* 在各语言专用字段中,以形态素(词素)为单位计数,被去除的助词等也会被计入。 + +因此,请用空格分隔单词,并将数字指定得宽松一些。 + +:: + + "全文 检索"~5 + "大阪 好吃"~10 + +不含空格的字符串(例如 ``"大阪好吃"~10``)在通用字段中会作为一个连续的字符串进行匹配,因此可能无法按预期作为单词之间的邻近检索工作。请用空格分隔单词。距离无法精确换算为字符数,因此建议先指定较大的数字,再根据检索结果逐步缩小。 + +相关主题 +---------- + +- :doc:`search-fuzzy` - 模糊检索 +- :doc:`search-boost` - 提升检索 +- :doc:`search-field` - 字段指定检索 +- :doc:`special-char` - 特殊字符与转义 diff --git a/zh-cn/15.9/user/special-char.rst b/zh-cn/15.9/user/special-char.rst index a3bb9110..0c9c2eba 100644 --- a/zh-cn/15.9/user/special-char.rst +++ b/zh-cn/15.9/user/special-char.rst @@ -8,7 +8,7 @@ + - && || ! ( ) { } [ ] ^ " ~ * ? : \ / -这些字符用于调用必需词与排除词(``+`` ``-``)、布尔运算符(``&&`` ``||`` ``!``)、分组(``( )``)、范围搜索(``[ ]`` ``{ }``)、权重搜索(``^``)、短语搜索(``"``)、模糊搜索(``~``)、通配符搜索(``*`` ``?``)以及字段搜索(``:``)等搜索功能。 +这些字符用于调用必需词与排除词(``+`` ``-``)、布尔运算符(``&&`` ``||`` ``!``)、分组(``( )``)、范围搜索(``[ ]`` ``{ }``)、权重搜索(``^``)、短语搜索(``"``)、模糊搜索与邻近搜索(``~``)、通配符搜索(``*`` ``?``)以及字段搜索(``:``)等搜索功能。 例如,搜索词中包含 URL 或文件路径里的"/"、":",或程序代码中的"+"、"-"等符号时,如果直接搜索而不加转义,可能会得到意料之外的搜索结果。转义方法请参见下文。 @@ -43,8 +43,8 @@ - 短语检索(将引号内的内容作为一个整体检索词处理,也可代替转义使用) - :doc:`advanced-search` * - ``~`` - - 模糊检索 - - :doc:`search-fuzzy` + - 模糊检索(用于单词后)/ 邻近检索(用于短语后) + - :doc:`search-fuzzy` / :doc:`search-proximity` * - ``*`` ``?`` - 通配符检索 - :doc:`search-wildcard` @@ -85,4 +85,5 @@ - :doc:`search-field` - 字段指定检索 - :doc:`search-wildcard` - 通配符检索 - :doc:`search-fuzzy` - 模糊检索 +- :doc:`search-proximity` - 邻近检索 - :doc:`advanced-search` - 高级检索