From c32187d7121ae2b0d510cb79f3025d5c3c3f6641 Mon Sep 17 00:00:00 2001 From: Shinsuke Sugaya Date: Sat, 26 Sep 2026 05:35:43 +0900 Subject: [PATCH 1/2] docs(15.9): list every log file in the English log file guide The English admin/log-guide.rst listed only six log files. The other six languages also describe fess-urls.log, searchlog.log and gc-crawler.log, which Fess writes to the logs directory and offers for download on the Log Files page. Add the three entries to the English page. --- en/15.9/admin/log-guide.rst | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/en/15.9/admin/log-guide.rst b/en/15.9/admin/log-guide.rst index 2324bc339..256f08840 100644 --- a/en/15.9/admin/log-guide.rst +++ b/en/15.9/admin/log-guide.rst @@ -47,4 +47,19 @@ fess-suggest.log Suggest log file. +fess-urls.log +::::::::::::: + +Crawl statistics for each URL, such as the time each processing step took. + +searchlog.log +::::::::::::: + +Search log file. + +gc-crawler.log +:::::::::::::: + +Garbage collection log file of the crawler. + .. |image0| image:: ../../../resources/images/en/15.9/admin/log-1.png From 4451dfacd3c79d02820a737a07697b963142fd99 Mon Sep 17 00:00:00 2001 From: Shinsuke Sugaya Date: Sat, 26 Sep 2026 05:35:44 +0900 Subject: [PATCH 2/2] docs(15.9): note that stopwords are matched per token A stopword is compared with each token the analyzer produces, not with the text as written. A word that the analyzer splits into several tokens, for example a word mixing letters and digits in Japanese analysis, is therefore not removed when it is added to the dictionary as it is, and following the language note on the page does not stop it from matching. Say so in the note and point to the _analyze API to check how a word is split. All seven languages, 15.9 tree only. --- de/15.9/admin/stopwords-guide.rst | 5 +++++ en/15.9/admin/stopwords-guide.rst | 4 ++++ es/15.9/admin/stopwords-guide.rst | 4 ++++ fr/15.9/admin/stopwords-guide.rst | 4 ++++ ja/15.9/admin/stopwords-guide.rst | 4 ++++ ko/15.9/admin/stopwords-guide.rst | 4 ++++ zh-cn/15.9/admin/stopwords-guide.rst | 4 ++++ 7 files changed, 29 insertions(+) diff --git a/de/15.9/admin/stopwords-guide.rst b/de/15.9/admin/stopwords-guide.rst index ff9a2fe1e..ce1804108 100644 --- a/de/15.9/admin/stopwords-guide.rst +++ b/de/15.9/admin/stopwords-guide.rst @@ -19,6 +19,11 @@ Im Stoppwort-Wörterbuch können Sie die zu entfernenden Wörter verwalten. wurde, kann daher über ein sprachspezifisches Feld weiterhin gefunden werden. Damit ein Wort nicht mehr gefunden wird, fügen Sie es auch dem Stoppwort-Wörterbuch der Sprache der Dokumente hinzu. + Stoppwörter werden mit jedem Token verglichen, das der Analyzer erzeugt. Ein Wort, das der Analyzer + in mehrere Token zerlegt, etwa ein Wort aus Buchstaben und Ziffern, wird nicht entfernt, wenn es + unverändert hinzugefügt wird. Wie ein Wort zerlegt wird, können Sie mit der ``_analyze``-API von + OpenSearch prüfen. + Verwaltung ========== diff --git a/en/15.9/admin/stopwords-guide.rst b/en/15.9/admin/stopwords-guide.rst index 872122384..36448e6ee 100644 --- a/en/15.9/admin/stopwords-guide.rst +++ b/en/15.9/admin/stopwords-guide.rst @@ -19,6 +19,10 @@ The stopwords dictionary lets you manage the words to remove. through a language-specific field. To keep a word from matching, also add it to the stopwords dictionary for the language of the documents. + Stopwords are compared with each token that the analyzer produces. A word that the analyzer + splits into several tokens, such as a word that mixes letters and digits, is not removed when it + is added as it is. You can check how a word is split with the ``_analyze`` API of OpenSearch. + Management Operations ===================== diff --git a/es/15.9/admin/stopwords-guide.rst b/es/15.9/admin/stopwords-guide.rst index 8378e07d8..17055be07 100644 --- a/es/15.9/admin/stopwords-guide.rst +++ b/es/15.9/admin/stopwords-guide.rst @@ -20,6 +20,10 @@ búsquedas. El diccionario de palabras vacías le permite administrar las palabr una palabra deje de coincidir, añádala también al diccionario de palabras vacías del idioma de los documentos. + Las palabras vacías se comparan con cada token que genera el analizador. Una palabra que el + analizador divide en varios tokens, como una palabra que mezcla letras y dígitos, no se elimina si + se añade tal cual. Puede comprobar cómo se divide una palabra con la API ``_analyze`` de OpenSearch. + Método de gestión ================== diff --git a/fr/15.9/admin/stopwords-guide.rst b/fr/15.9/admin/stopwords-guide.rst index 39e86a1c4..9cdeb6afa 100644 --- a/fr/15.9/admin/stopwords-guide.rst +++ b/fr/15.9/admin/stopwords-guide.rst @@ -20,6 +20,10 @@ recherches. Le dictionnaire de mots vides permet de gérer les mots à supprimer langue. Pour qu'un mot ne corresponde plus, ajoutez-le aussi au dictionnaire de mots vides de la langue des documents. + Les mots vides sont comparés à chaque jeton produit par l'analyseur. Un mot que l'analyseur découpe + en plusieurs jetons, comme un mot mêlant lettres et chiffres, n'est pas supprimé s'il est ajouté tel + quel. Vous pouvez vérifier le découpage d'un mot avec l'API ``_analyze`` d'OpenSearch. + Gestion ======= diff --git a/ja/15.9/admin/stopwords-guide.rst b/ja/15.9/admin/stopwords-guide.rst index 739217f80..51f901764 100644 --- a/ja/15.9/admin/stopwords-guide.rst +++ b/ja/15.9/admin/stopwords-guide.rst @@ -18,6 +18,10 @@ ``en/stopwords.txt`` にだけ追加した単語は、言語別フィールドを通じて検索にヒットし続けることがあります。 検索にヒットさせたくない単語は、対象の文書の言語のストップワード辞書にも追加してください。 + ストップワードはアナライザーが分割した後の単語(トークン)ごとに照合されます。 + 英字と数字が混ざった語など、アナライザーが複数のトークンに分割する語は、その語のまま追加しても取り除かれません。 + 分割のされ方は OpenSearch の ``_analyze`` API で確認できます。 + 管理方法 ====== diff --git a/ko/15.9/admin/stopwords-guide.rst b/ko/15.9/admin/stopwords-guide.rst index 3eb114355..a76f3a428 100644 --- a/ko/15.9/admin/stopwords-guide.rst +++ b/ko/15.9/admin/stopwords-guide.rst @@ -18,6 +18,10 @@ ``en/stopwords.txt`` 에만 추가한 단어는 언어별 필드를 통해 계속 검색될 수 있습니다. 검색되지 않게 하려는 단어는 대상 문서 언어의 불용어 사전에도 추가하십시오. + 불용어는 애널라이저가 분할한 단어(토큰)마다 비교됩니다. + 영문자와 숫자가 섞인 단어처럼 애널라이저가 여러 토큰으로 분할하는 단어는 그대로 추가해도 제거되지 않습니다. + 단어가 어떻게 분할되는지는 OpenSearch의 ``_analyze`` API로 확인할 수 있습니다. + 관리 방법 ====== diff --git a/zh-cn/15.9/admin/stopwords-guide.rst b/zh-cn/15.9/admin/stopwords-guide.rst index 6a11932c9..d3f84e119 100644 --- a/zh-cn/15.9/admin/stopwords-guide.rst +++ b/zh-cn/15.9/admin/stopwords-guide.rst @@ -18,6 +18,10 @@ 仍可能通过语言字段被搜索命中。 对于不希望被搜索命中的词,请同时将其添加到目标文档语言的停用词词典中。 + 停用词按分析器切分后的每个词(词元)进行比对。 + 像字母与数字混合的词这样会被分析器切分为多个词元的词,即使原样添加也不会被去除。 + 可以使用 OpenSearch 的 ``_analyze`` API 确认词的切分方式。 + 管理方法 ======