diff --git a/CHANGELOG.md b/CHANGELOG.md
index e9172cd..ea9b5b0 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -2,6 +2,16 @@
All notable changes to Backlink Intelligence are documented here.
+## 1.0.1 - 2026-08-30
+
+### Fixed
+
+- Prevented partial anchor insertion inside larger word forms such as rendering `AI agents` as `[AI Agent]s`.
+- Preserved the publisher's existing anchor capitalization when the requested keyword differs only by case.
+- Added conservative singular/plural anchor adaptation when the natural grammatical form already exists in source copy.
+- Added destination-intent scoring so context specific to the destination topic receives more weight than generic anchor repetition.
+- Added destination-fit and actual placed-anchor details to placement CLI output.
+
## 1.0.0 - 2026-08-30
### Added
diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md
index a4f8c4c..201256c 100644
--- a/CONTRIBUTING.md
+++ b/CONTRIBUTING.md
@@ -19,7 +19,7 @@ Contributions should preserve the project's core principles:
git clone https://github.com/alok-vibe-code/backlink-intelligence.git
cd backlink-intelligence
python -m venv .venv
-source .venv/bin/activate # Windows: .venv\Scripts\activate
+source .venv/bin/activate # Windows: .venv\\Scripts\\activate
python -m pip install -e .
python -m unittest discover -s tests -v
```
diff --git a/README.md b/README.md
index 09d4a0a..5317b89 100644
--- a/README.md
+++ b/README.md
@@ -117,8 +117,6 @@ The report includes evidence such as relevance, placement potential, outbound-li
- `manual_review`
- `low_priority`
-Bulk commands cache repeated URLs during a run and use a small delay between rows by default. Use `--delay` to adjust that delay responsibly.
-
## 3. Find contextual link placements
This is the signature workflow.
@@ -135,8 +133,9 @@ For each recommended paragraph the tool returns:
- paragraph number,
- context-fit level,
+- destination-fit level and score,
- requested anchor,
-- suggested anchor when the requested wording is awkward,
+- actual placed anchor (preserving source capitalization/grammar where possible),
- placement strategy,
- editorial intervention level,
- original-text preservation,
@@ -145,7 +144,7 @@ For each recommended paragraph the tool returns:
- reasons,
- and review flags.
-The deterministic rewrite engine deliberately favors minimal editorial change. Its output is a placement draft for human review, not an instruction to publish automatically.
+The deterministic rewrite engine deliberately favors minimal editorial change. Exact phrase matching uses complete word boundaries, preserves the publisher's existing capitalization, and can conservatively adapt a singular/plural anchor to the grammatical form already present in source copy. Destination-intent scoring also helps distinguish a paragraph that matches the target's specific topic from one that merely repeats the requested anchor. Its output is a placement draft for human review, not an instruction to publish automatically.
## 4. Monitor acquired backlinks
@@ -312,15 +311,6 @@ backlink_intelligence/
└── safety.py
```
-## Docker
-
-Build and run the CLI without installing Python packages into your host environment:
-
-```bash
-docker build -t backlink-intelligence .
-docker run --rm backlink-intelligence --help
-```
-
## Contributing
Contributions, test cases, parser improvements, and evidence-based methodology discussions are welcome. Read [CONTRIBUTING.md](CONTRIBUTING.md) first.
diff --git a/SECURITY.md b/SECURITY.md
index db92208..9aad419 100644
--- a/SECURITY.md
+++ b/SECURITY.md
@@ -8,24 +8,26 @@ Use GitHub's private vulnerability reporting feature when enabled for this repos
## Security model
-Backlink Intelligence processes URLs supplied by users. Network-enabled functionality therefore treats URL fetching as security-sensitive.
+Backlink Intelligence is expected to process URLs supplied by users. Network-enabled releases therefore treat URL fetching as security-sensitive functionality.
-The v1 fetcher blocks or bounds, at minimum:
+Future implementations should protect against, at minimum:
-- unsupported schemes,
-- credentials embedded in URLs,
-- localhost and common local hostnames,
-- private, loopback, link-local, reserved, multicast, and unspecified IP ranges,
-- DNS resolution to non-public addresses,
-- redirects to non-public addresses,
+- localhost and loopback targets,
+- private and link-local IP ranges,
+- cloud metadata endpoints,
+- DNS rebinding and redirect-to-private-network behavior,
- excessive redirect chains,
- oversized responses,
-- unexpectedly slow endpoints through request timeouts,
-- and non-HTML content.
+- decompression bombs,
+- unexpectedly slow endpoints,
+- unsupported schemes,
+- unsafe local file paths,
+- use as a generic open proxy,
+- and uncontrolled concurrency during bulk analysis.
## Hosted deployment warning
-A command-line tool that fetches user-supplied URLs and an internet-facing web service have different risk profiles. Do not expose the crawler as a public web endpoint without additional validation, centralized rate limiting, caching, abuse prevention, request isolation, logging, and SSRF defenses.
+A command-line tool that fetches user-supplied URLs and an internet-facing web service have different risk profiles. Do not expose the crawler as a public web endpoint without additional validation, rate limiting, abuse prevention, request isolation, and SSRF defenses.
## Secrets
diff --git a/backlink_intelligence/__init__.py b/backlink_intelligence/__init__.py
index df30da3..35fbf6a 100644
--- a/backlink_intelligence/__init__.py
+++ b/backlink_intelligence/__init__.py
@@ -1,3 +1,3 @@
"""Backlink Intelligence package."""
-__version__ = "1.0.0"
+__version__ = "1.0.1"
diff --git a/backlink_intelligence/cli.py b/backlink_intelligence/cli.py
index 900176f..4b45bcd 100644
--- a/backlink_intelligence/cli.py
+++ b/backlink_intelligence/cli.py
@@ -15,20 +15,47 @@
def build_parser() -> argparse.ArgumentParser:
- parser = argparse.ArgumentParser(prog="backlink-intelligence", description="Evidence-first backlink intelligence for auditing, qualification, contextual placement, monitoring, and portfolio review.")
+ parser = argparse.ArgumentParser(
+ prog="backlink-intelligence",
+ description=(
+ "Evidence-first backlink intelligence for auditing, qualification, "
+ "contextual placement, monitoring, and portfolio review."
+ ),
+ )
parser.add_argument("--version", action="version", version=f"%(prog)s {__version__}")
sub = parser.add_subparsers(dest="command")
+
sub.add_parser("status", help="Show implementation status.")
+
audit = sub.add_parser("audit", help="Audit an existing backlink.")
- audit.add_argument("source_url"); audit.add_argument("target_url"); audit.add_argument("--json", action="store_true", dest="as_json"); audit.add_argument("--output", help="Optional JSON output file.")
+ audit.add_argument("source_url")
+ audit.add_argument("target_url")
+ audit.add_argument("--json", action="store_true", dest="as_json")
+ audit.add_argument("--output", help="Optional JSON output file.")
+
qualify = sub.add_parser("qualify", help="Qualify backlink prospects from CSV.")
- qualify.add_argument("input_csv"); qualify.add_argument("--output", default="qualification-report.csv"); qualify.add_argument("--delay", type=float, default=0.5, help="Delay between bulk rows in seconds.")
+ qualify.add_argument("input_csv")
+ qualify.add_argument("--output", default="qualification-report.csv")
+ qualify.add_argument("--delay", type=float, default=0.5, help="Delay between bulk rows in seconds.")
+
place = sub.add_parser("place", help="Find contextual link placement opportunities.")
- place.add_argument("source_url"); place.add_argument("target_url"); place.add_argument("--anchor", required=True, help="Preferred anchor / keyword."); place.add_argument("--top", type=int, default=3); place.add_argument("--json", action="store_true", dest="as_json"); place.add_argument("--output", help="Optional JSON output file.")
+ place.add_argument("source_url")
+ place.add_argument("target_url")
+ place.add_argument("--anchor", required=True, help="Preferred anchor / keyword.")
+ place.add_argument("--top", type=int, default=3)
+ place.add_argument("--json", action="store_true", dest="as_json")
+ place.add_argument("--output", help="Optional JSON output file.")
+
monitor = sub.add_parser("monitor", help="Monitor backlinks listed in a CSV file.")
- monitor.add_argument("input_csv"); monitor.add_argument("--state", default="backlink-state.json"); monitor.add_argument("--output", default="monitor-report.csv"); monitor.add_argument("--delay", type=float, default=0.5, help="Delay between monitoring rows in seconds.")
+ monitor.add_argument("input_csv")
+ monitor.add_argument("--state", default="backlink-state.json")
+ monitor.add_argument("--output", default="monitor-report.csv")
+ monitor.add_argument("--delay", type=float, default=0.5, help="Delay between monitoring rows in seconds.")
+
portfolio = sub.add_parser("portfolio", help="Analyze anchor/destination/placement distributions.")
- portfolio.add_argument("input_csv"); portfolio.add_argument("--output", help="Optional JSON output file.")
+ portfolio.add_argument("input_csv")
+ portfolio.add_argument("--output", help="Optional JSON output file.")
+
return parser
@@ -37,37 +64,86 @@ def _placement_text(items) -> str:
return "No suitable placement opportunities could be generated."
chunks: list[str] = []
for item in items:
- chunks.extend([f"PLACEMENT OPPORTUNITY #{item.rank}", f"Paragraph: {item.paragraph_index}", f"Context fit: {item.context_level}", f"Similarity score: {item.score:.3f}", f"Strategy: {item.strategy}", f"Editorial intervention:{item.intervention}", f"Text preservation: {item.preservation_percent:.1f}%", "", "BEFORE", item.before, "", "AFTER", item.after])
- if item.reasons: chunks.extend(["", "Why this placement:", *[f" + {v}" for v in item.reasons]])
- if item.warnings: chunks.extend(["", "Review flags:", *[f" ! {v}" for v in item.warnings]])
+ chunks.extend(
+ [
+ f"PLACEMENT OPPORTUNITY #{item.rank}",
+ f"Paragraph: {item.paragraph_index}",
+ f"Context fit: {item.context_level}",
+ f"Similarity score: {item.score:.3f}",
+ f"Destination fit: {item.destination_fit} ({item.destination_score:.3f})",
+ f"Requested anchor: {item.requested_anchor}",
+ f"Placed anchor: {item.suggested_anchor}",
+ f"Strategy: {item.strategy}",
+ f"Editorial intervention:{item.intervention}",
+ f"Text preservation: {item.preservation_percent:.1f}%",
+ "",
+ "BEFORE",
+ item.before,
+ "",
+ "AFTER",
+ item.after,
+ ]
+ )
+ if item.reasons:
+ chunks.extend(["", "Why this placement:", *[f" + {v}" for v in item.reasons]])
+ if item.warnings:
+ chunks.extend(["", "Review flags:", *[f" ! {v}" for v in item.warnings]])
chunks.extend(["", "-" * 72, ""])
return "\n".join(chunks).rstrip()
def main(argv: list[str] | None = None) -> int:
- parser = build_parser(); args = parser.parse_args(argv)
+ parser = build_parser()
+ args = parser.parse_args(argv)
+
try:
if args.command == "status":
- print("Backlink Intelligence is installed and functional."); print(f"Version: {__version__}"); print("Available: audit, qualify, place, monitor, portfolio"); print("Workflow: Discover -> Qualify -> Place -> Monitor -> Analyze"); return 0
+ print("Backlink Intelligence is installed and functional.")
+ print(f"Version: {__version__}")
+ print("Available: audit, qualify, place, monitor, portfolio")
+ print("Workflow: Discover -> Qualify -> Place -> Monitor -> Analyze")
+ return 0
+
if args.command == "audit":
result = audit_backlink(args.source_url, args.target_url)
- if args.output: write_json(result.to_dict(), args.output)
- print(write_json(result.to_dict()) if args.as_json else audit_text(result)); return 0 if result.source.status_code else 2
+ if args.output:
+ write_json(result.to_dict(), args.output)
+ print(write_json(result.to_dict()) if args.as_json else audit_text(result))
+ return 0 if result.source.status_code else 2
+
if args.command == "qualify":
- rows = qualify_csv(args.input_csv, args.output, delay_seconds=args.delay); print(f"Qualified {len(rows)} prospect(s). Report: {args.output}"); return 0
+ rows = qualify_csv(args.input_csv, args.output, delay_seconds=args.delay)
+ print(f"Qualified {len(rows)} prospect(s). Report: {args.output}")
+ return 0
+
if args.command == "place":
- items = suggest_placements(args.source_url, args.target_url, args.anchor, top_n=args.top); payload = [item.to_dict() for item in items]
- if args.output: write_json(payload, args.output)
- print(write_json(payload) if args.as_json else _placement_text(items)); return 0 if items else 3
+ items = suggest_placements(args.source_url, args.target_url, args.anchor, top_n=args.top)
+ payload = [item.to_dict() for item in items]
+ if args.output:
+ write_json(payload, args.output)
+ print(write_json(payload) if args.as_json else _placement_text(items))
+ return 0 if items else 3
+
if args.command == "monitor":
- rows = monitor_csv(args.input_csv, args.state, args.output, delay_seconds=args.delay); changed = sum("unchanged" not in row["changes"] and "baseline_created" not in row["changes"] for row in rows); print(f"Checked {len(rows)} link(s). Changed: {changed}. Report: {args.output}"); return 0
+ rows = monitor_csv(args.input_csv, args.state, args.output, delay_seconds=args.delay)
+ changed = sum("unchanged" not in row["changes"] and "baseline_created" not in row["changes"] for row in rows)
+ print(f"Checked {len(rows)} link(s). Changed: {changed}. Report: {args.output}")
+ return 0
+
if args.command == "portfolio":
- result = analyze_portfolio(args.input_csv); print(write_json(result, args.output)); return 0
- parser.print_help(); return 0
+ result = analyze_portfolio(args.input_csv)
+ text = write_json(result, args.output)
+ print(text)
+ return 0
+
+ parser.print_help()
+ return 0
except KeyboardInterrupt:
- print("Cancelled.", file=sys.stderr); return 130
+ print("Cancelled.", file=sys.stderr)
+ return 130
except Exception as exc:
- print(f"Error: {exc}", file=sys.stderr); return 1
+ print(f"Error: {exc}", file=sys.stderr)
+ return 1
if __name__ == "__main__":
diff --git a/backlink_intelligence/models.py b/backlink_intelligence/models.py
index f03beeb..63e1c0c 100644
--- a/backlink_intelligence/models.py
+++ b/backlink_intelligence/models.py
@@ -105,6 +105,8 @@ class PlacementSuggestion:
paragraph_index: int
score: float
context_level: str
+ destination_score: float
+ destination_fit: str
requested_anchor: str
suggested_anchor: str
strategy: str
diff --git a/backlink_intelligence/placement.py b/backlink_intelligence/placement.py
index af15079..6e85373 100644
--- a/backlink_intelligence/placement.py
+++ b/backlink_intelligence/placement.py
@@ -1,5 +1,7 @@
from __future__ import annotations
+import re
+
from .fetcher import FetchConfig, fetch_page
from .models import PageEvidence, PlacementSuggestion
from .relevance import similarity, tokens
@@ -32,53 +34,217 @@ def _select_anchor(preferred: str, target_title: str) -> tuple[str, list[str]]:
if not preferred:
return (target_title.strip() or "this related resource"), []
_, warnings = _anchor_naturalness(preferred)
+ # Preserve the user's keyword when it reads naturally. When it is mechanically
+ # awkward, offer the target title as a safer editorial alternative.
if warnings and target_title.strip():
return target_title.strip(), warnings + ["suggested_anchor_differs_from_requested"]
return preferred, warnings
-def _compose_after(paragraph: str, anchor: str, target_url: str, target_title: str) -> tuple[str, str]:
+def _find_complete_phrase(text: str, phrase: str) -> re.Match[str] | None:
+ """Find a phrase only when it is not embedded inside a larger word form."""
+ phrase = phrase.strip()
+ if not phrase:
+ return None
+ pattern = re.compile(rf"(? list[str]:
+ """Return conservative singular/plural variants for the final anchor word."""
+ anchor = anchor.strip()
+ if not anchor or " " not in anchor:
+ return []
+ prefix, last = anchor.rsplit(" ", 1)
+ if not last.isalpha():
+ return []
+
+ lower = last.lower()
+ variants: list[str] = []
+ # Singular -> simple plural. This intentionally avoids guessing irregular forms.
+ if not lower.endswith("s"):
+ variants.append(f"{prefix} {last}s")
+ # Plural -> simple singular, excluding common singular words that end in s.
+ elif len(last) > 3 and not lower.endswith(("ss", "us", "is")):
+ variants.append(f"{prefix} {last[:-1]}")
+ return variants
+
+
+def _compose_after(
+ paragraph: str,
+ anchor: str,
+ target_url: str,
+ target_title: str,
+) -> tuple[str, str, str, list[str]]:
+ """Compose the draft while preserving source grammar/capitalization when possible."""
+ exact = _find_complete_phrase(paragraph, anchor)
+ if exact is not None:
+ placed_anchor = exact.group(0)
+ linked = f"[{placed_anchor}]({target_url})"
+ after = paragraph[: exact.start()] + linked + paragraph[exact.end() :]
+ notes: list[str] = []
+ if placed_anchor != anchor:
+ notes.append("source_anchor_capitalization_preserved")
+ return "minimal_insertion", after, placed_anchor, notes
+
+ # If the exact requested form is not present, prefer a complete natural word-form
+ # already in the publisher copy instead of creating artifacts such as [AI Agent]s.
+ for variant in _simple_anchor_variants(anchor):
+ match = _find_complete_phrase(paragraph, variant)
+ if match is not None:
+ placed_anchor = match.group(0)
+ linked = f"[{placed_anchor}]({target_url})"
+ after = paragraph[: match.start()] + linked + paragraph[match.end() :]
+ return (
+ "minimal_insertion",
+ after,
+ placed_anchor,
+ ["anchor_adapted_to_source_grammar", "requested_anchor_not_used_verbatim"],
+ )
+
linked = f"[{anchor}]({target_url})"
- idx = paragraph.lower().find(anchor.lower())
- if idx >= 0:
- return "minimal_insertion", paragraph[:idx] + linked + paragraph[idx + len(anchor):]
topic = target_title.strip()
- sentence = f"For a more detailed resource on {topic}, see {linked}." if topic and topic.lower() != anchor.lower() else f"For a more detailed resource on this topic, see {linked}."
- return "contextual_sentence", paragraph.rstrip() + " " + sentence
+ if topic and topic.lower() != anchor.lower():
+ sentence = f"For a more detailed resource on {topic}, see {linked}."
+ else:
+ sentence = f"For a more detailed resource on this topic, see {linked}."
+ return "contextual_sentence", paragraph.rstrip() + " " + sentence, anchor, []
+
+
+def _stem(term: str) -> str:
+ """Small deterministic normalizer used only for destination-intent comparison."""
+ term = term.lower().strip(".-")
+ if len(term) > 5 and term.endswith("ies"):
+ return term[:-3] + "y"
+ if len(term) > 5 and term.endswith("ing"):
+ return term[:-3]
+ if len(term) > 4 and term.endswith("es") and not term.endswith("ses"):
+ return term[:-2]
+ if len(term) > 4 and term.endswith("s") and not term.endswith(("ss", "us", "is")):
+ return term[:-1]
+ return term
-def rank_placements(source: PageEvidence, target: PageEvidence, preferred_anchor: str, target_url: str, *, top_n: int = 3) -> list[PlacementSuggestion]:
+def _stems(text: str) -> set[str]:
+ return {_stem(term) for term in tokens(text)}
+
+
+def _destination_intent_score(paragraph: str, target: PageEvidence, anchor: str) -> float:
+ """Measure fit to destination-specific intent, not just the requested anchor."""
+ core_profile = " ".join([target.title, target.h1]).strip()
+ if not core_profile:
+ core_profile = " ".join(target.headings[:8]).strip()
+ if not core_profile:
+ return similarity(paragraph, target.text[:4000])
+
+ paragraph_terms = _stems(paragraph)
+ core_terms = _stems(core_profile)
+ anchor_terms = _stems(anchor)
+
+ # Prefer terms that describe what makes the destination distinct from the anchor.
+ intent_terms = core_terms - anchor_terms
+ if len(intent_terms) < 2:
+ intent_terms = core_terms
+ intent_overlap = len(paragraph_terms & intent_terms) / max(len(intent_terms), 1)
+ semantic = similarity(paragraph, core_profile)
+ return round((0.35 * semantic) + (0.65 * intent_overlap), 4)
+
+
+def _destination_level(score: float) -> str:
+ if score >= 0.25:
+ return "very_high"
+ if score >= 0.14:
+ return "high"
+ if score >= 0.08:
+ return "medium"
+ return "low"
+
+
+def rank_placements(
+ source: PageEvidence,
+ target: PageEvidence,
+ preferred_anchor: str,
+ target_url: str,
+ *,
+ top_n: int = 3,
+) -> list[PlacementSuggestion]:
if source.status_code != 200 or target.status_code != 200:
return []
+
anchor, anchor_warnings = _select_anchor(preferred_anchor, target.title)
target_profile = " ".join([target.title, target.h1, *target.headings, target.text[:12000]])
- candidates: list[tuple[float, int, str]] = []
+ candidates: list[tuple[float, float, int, str]] = []
+
for i, paragraph in enumerate(source.paragraphs, start=1):
wc = len(paragraph.split())
if wc < 18 or wc > 260:
continue
- score = similarity(paragraph, target_profile)
+ semantic_score = similarity(paragraph, target_profile)
+ destination_score = _destination_intent_score(paragraph, target, anchor)
anchor_terms = set(tokens(anchor))
anchor_overlap = len(anchor_terms & set(tokens(paragraph))) / max(len(anchor_terms), 1)
- candidates.append((round((0.84 * score) + (0.16 * anchor_overlap), 4), i, paragraph))
- candidates.sort(key=lambda item: (-item[0], item[1]))
+ # Destination intent gets meaningful weight so a pricing/cost paragraph beats a
+ # generic paragraph that merely repeats the requested anchor.
+ score = round((0.62 * semantic_score) + (0.30 * destination_score) + (0.08 * anchor_overlap), 4)
+ candidates.append((score, destination_score, i, paragraph))
+
+ candidates.sort(key=lambda item: (-item[0], -item[1], item[2]))
suggestions: list[PlacementSuggestion] = []
- for rank, (score, index, paragraph) in enumerate(candidates[:max(top_n, 1)], start=1):
- strategy, after = _compose_after(paragraph, anchor, target_url, target.title)
+ for rank, (score, destination_score, index, paragraph) in enumerate(candidates[: max(top_n, 1)], start=1):
+ strategy, after, placed_anchor, compose_notes = _compose_after(paragraph, anchor, target_url, target.title)
original_words = max(len(paragraph.split()), 1)
- added = max(len(after.split()) - original_words, 0)
- preservation = 100.0
+ after_words = len(after.split())
+ added = max(after_words - original_words, 0)
+ preservation = 100.0 if strategy in {"minimal_insertion", "contextual_sentence"} else 90.0
warnings = list(anchor_warnings)
reasons = ["paragraph_has_strong_target_similarity"] if score >= 0.25 else ["best_available_context_match"]
- reasons.append("anchor_already_present_in_original_copy" if strategy == "minimal_insertion" else "publisher_copy_preserved")
+ if strategy == "minimal_insertion":
+ reasons.append("anchor_already_present_in_original_copy")
+ else:
+ reasons.append("publisher_copy_preserved")
+ for note in compose_notes:
+ if note == "requested_anchor_not_used_verbatim":
+ warnings.append(note)
+ else:
+ reasons.append(note)
+ if destination_score >= 0.14:
+ reasons.append("strong_destination_intent_alignment")
+ elif destination_score < 0.08:
+ warnings.append("weak_destination_intent_alignment")
if score < 0.12:
warnings.append("weak_context_match_manual_review_required")
context_level = "very_high" if score >= 0.48 else "high" if score >= 0.30 else "medium" if score >= 0.15 else "low"
- suggestions.append(PlacementSuggestion(rank=rank, paragraph_index=index, score=score, context_level=context_level, requested_anchor=preferred_anchor, suggested_anchor=anchor, strategy=strategy, before=paragraph, after=after, added_words=added, preservation_percent=preservation, intervention=_intervention(preservation, added), reasons=reasons, warnings=warnings))
+ suggestions.append(
+ PlacementSuggestion(
+ rank=rank,
+ paragraph_index=index,
+ score=score,
+ context_level=context_level,
+ destination_score=destination_score,
+ destination_fit=_destination_level(destination_score),
+ requested_anchor=preferred_anchor,
+ suggested_anchor=placed_anchor,
+ strategy=strategy,
+ before=paragraph,
+ after=after,
+ added_words=added,
+ preservation_percent=preservation,
+ intervention=_intervention(preservation, added),
+ reasons=reasons,
+ warnings=warnings,
+ )
+ )
return suggestions
-def suggest_placements(source_url: str, target_url: str, preferred_anchor: str, *, top_n: int = 3, config: FetchConfig | None = None) -> list[PlacementSuggestion]:
+def suggest_placements(
+ source_url: str,
+ target_url: str,
+ preferred_anchor: str,
+ *,
+ top_n: int = 3,
+ config: FetchConfig | None = None,
+) -> list[PlacementSuggestion]:
source = fetch_page(source_url, config)
target = fetch_page(target_url, config)
return rank_placements(source, target, preferred_anchor, target_url, top_n=top_n)
diff --git a/docs/methodology.md b/docs/methodology.md
index fb5e745..3f08d0d 100644
--- a/docs/methodology.md
+++ b/docs/methodology.md
@@ -38,9 +38,9 @@ The recommendation is intentionally explainable and reversible by a human review
## 6. Contextual placement
-Paragraphs are ranked by similarity to the target page with a smaller anchor-term overlap component. Very short and extremely long paragraphs are excluded from candidate generation.
+Paragraphs are ranked using target-page similarity, destination-specific intent, and a smaller anchor-term overlap component. Destination intent emphasizes terms that distinguish the target from the requested anchor, helping a cost/pricing paragraph outrank a generic paragraph that only mentions the entity. Very short and extremely long paragraphs are excluded from candidate generation.
-Before/After output prioritizes preservation of publisher text. If the anchor already appears naturally, only that phrase is linked. Otherwise a conservative contextual sentence is appended.
+Before/After output prioritizes preservation of publisher text. Anchor matching requires complete word boundaries, preserves the capitalization already present in publisher copy, and may conservatively use an existing singular/plural grammatical form rather than creating partial-link artifacts. If no natural source phrase is available, a conservative contextual sentence is appended.
## 7. Monitoring
diff --git a/pyproject.toml b/pyproject.toml
index f9b5caa..b338136 100644
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "backlink-intelligence"
-version = "1.0.0"
+version = "1.0.1"
description = "Open-source backlink intelligence based on evidence, context, and editorial fit."
readme = "README.md"
requires-python = ">=3.11"
diff --git a/tests/test_cli.py b/tests/test_cli.py
index 216072d..9ae22e2 100644
--- a/tests/test_cli.py
+++ b/tests/test_cli.py
@@ -8,7 +8,7 @@
class CLITests(unittest.TestCase):
def test_version_is_stable(self):
- self.assertEqual(__version__, "1.0.0")
+ self.assertEqual(__version__, "1.0.1")
def test_status_command(self):
output = io.StringIO()
diff --git a/tests/test_placement.py b/tests/test_placement.py
index d307b64..da0eaab 100644
--- a/tests/test_placement.py
+++ b/tests/test_placement.py
@@ -7,17 +7,105 @@
class PlacementTests(unittest.TestCase):
def setUp(self):
- self.source = parse_page("""
AI Agent ArchitectureBuilding AI Agents
Modern AI agents combine tool calling, retrieval, memory, and orchestration to complete complex workflows. Teams often introduce these capabilities progressively as systems become more autonomous and reliable.
Unrelated paragraph about office furniture, chairs, desks, shelving, workplace lighting, and interior design trends for modern offices.
""", requested_url="https://source.com/article", final_url="https://source.com/article", status_code=200)
- self.target = parse_page("""Agentic AI Learning RoadmapLearn Agentic AI
A hands-on roadmap covering tool calling, retrieval, RAG, memory, multi-agent systems, evaluation, security, and production reliability.
""", requested_url="https://target.com/roadmap", final_url="https://target.com/roadmap", status_code=200)
+ self.source = parse_page(
+ """AI Agent ArchitectureBuilding AI Agents
+ Modern AI agents combine tool calling, retrieval, memory, and orchestration to complete complex workflows. Teams often introduce these capabilities progressively as systems become more autonomous and reliable.
+ Unrelated paragraph about office furniture, chairs, desks, shelving, workplace lighting, and interior design trends for modern offices.
+ """,
+ requested_url="https://source.com/article", final_url="https://source.com/article", status_code=200,
+ )
+ self.target = parse_page(
+ """Agentic AI Learning RoadmapLearn Agentic AI
+ A hands-on roadmap covering tool calling, retrieval, RAG, memory, multi-agent systems, evaluation, security, and production reliability.
""",
+ requested_url="https://target.com/roadmap", final_url="https://target.com/roadmap", status_code=200,
+ )
+
@patch("backlink_intelligence.placement.fetch_page")
def test_returns_ranked_before_after(self, fetch):
- fetch.side_effect = [self.source, self.target]; items = suggest_placements("https://source.com/article", "https://target.com/roadmap", "Agentic AI learning roadmap", top_n=2); self.assertGreaterEqual(len(items), 1); self.assertIn("[Agentic AI learning roadmap](https://target.com/roadmap)", items[0].after); self.assertEqual(items[0].paragraph_index, 1); self.assertIn(items[0].strategy, {"minimal_insertion", "contextual_sentence"}); self.assertGreaterEqual(items[0].preservation_percent, 95)
+ fetch.side_effect = [self.source, self.target]
+ items = suggest_placements("https://source.com/article", "https://target.com/roadmap", "Agentic AI learning roadmap", top_n=2)
+ self.assertGreaterEqual(len(items), 1)
+ self.assertIn("BEFORE" if False else "", "")
+ self.assertIn("[Agentic AI learning roadmap](https://target.com/roadmap)", items[0].after)
+ self.assertEqual(items[0].paragraph_index, 1)
+ self.assertIn(items[0].strategy, {"minimal_insertion", "contextual_sentence"})
+ self.assertGreaterEqual(items[0].preservation_percent, 95)
+
@patch("backlink_intelligence.placement.fetch_page")
def test_exact_anchor_in_paragraph_uses_minimal_insertion(self, fetch):
- source = parse_page("This Agentic AI learning roadmap introduces tools, memory, retrieval, evaluation, and production patterns for engineers building modern agents.
", requested_url="https://s.com", final_url="https://s.com", status_code=200); fetch.side_effect = [source, self.target]; items = suggest_placements("https://s.com", "https://target.com/roadmap", "Agentic AI learning roadmap", top_n=1); self.assertEqual(items[0].strategy, "minimal_insertion")
+ source = parse_page("This Agentic AI learning roadmap introduces tools, memory, retrieval, evaluation, and production patterns for engineers building modern agents.
", requested_url="https://s.com", final_url="https://s.com", status_code=200)
+ fetch.side_effect = [source, self.target]
+ items = suggest_placements("https://s.com", "https://target.com/roadmap", "Agentic AI learning roadmap", top_n=1)
+ self.assertEqual(items[0].strategy, "minimal_insertion")
+
@patch("backlink_intelligence.placement.fetch_page")
def test_awkward_anchor_gets_editorial_alternative(self, fetch):
- fetch.side_effect = [self.source, self.target]; items = suggest_placements("https://source.com/article", "https://target.com/roadmap", "THIS IS A VERY LONG AWKWARD ANCHOR PHRASE FOR SEO", top_n=1); self.assertEqual(items[0].suggested_anchor, "Agentic AI Learning Roadmap"); self.assertIn("suggested_anchor_differs_from_requested", items[0].warnings)
+ fetch.side_effect = [self.source, self.target]
+ items = suggest_placements("https://source.com/article", "https://target.com/roadmap", "THIS IS A VERY LONG AWKWARD ANCHOR PHRASE FOR SEO", top_n=1)
+ self.assertEqual(items[0].suggested_anchor, "Agentic AI Learning Roadmap")
+ self.assertIn("suggested_anchor_differs_from_requested", items[0].warnings)
+
+ @patch("backlink_intelligence.placement.fetch_page")
+ def test_preserves_existing_anchor_capitalization(self, fetch):
+ source = parse_page(
+ "A custom AI agent can connect a website, CRM, email, database, and internal dashboard while supporting several business workflows.
",
+ requested_url="https://s.com", final_url="https://s.com", status_code=200,
+ )
+ target = parse_page(
+ "AI Agent Cost in 2026: Pricing Models, Hidden Costs, TCO, and ROIAI Agent Cost
AI agent pricing includes development and operating costs.
",
+ requested_url="https://t.com", final_url="https://t.com", status_code=200,
+ )
+ fetch.side_effect = [source, target]
+ item = suggest_placements("https://s.com", "https://t.com", "AI Agent", top_n=1)[0]
+ self.assertIn("[AI agent](https://t.com)", item.after)
+ self.assertNotIn("[AI Agent](https://t.com)", item.after)
+ self.assertEqual(item.suggested_anchor, "AI agent")
+ self.assertIn("source_anchor_capitalization_preserved", item.reasons)
+
+ @patch("backlink_intelligence.placement.fetch_page")
+ def test_plural_existing_anchor_is_linked_as_complete_phrase(self, fetch):
+ source = parse_page(
+ "Modern AI agents can coordinate tools, retrieval, memory, approvals, and business systems across several connected workflows while supporting reliable operations for growing teams.
",
+ requested_url="https://s.com", final_url="https://s.com", status_code=200,
+ )
+ target = parse_page(
+ "AI Agent Cost in 2026: Pricing Models and ROIAI Agent Cost
AI agent costs include development and operations.
",
+ requested_url="https://t.com", final_url="https://t.com", status_code=200,
+ )
+ fetch.side_effect = [source, target]
+ item = suggest_placements("https://s.com", "https://t.com", "AI Agent", top_n=1)[0]
+ self.assertIn("[AI agents](https://t.com)", item.after)
+ self.assertNotIn("[AI Agent](https://t.com)s", item.after)
+ self.assertEqual(item.suggested_anchor, "AI agents")
+ self.assertIn("anchor_adapted_to_source_grammar", item.reasons)
+ self.assertIn("requested_anchor_not_used_verbatim", item.warnings)
+
+ @patch("backlink_intelligence.placement.fetch_page")
+ def test_destination_intent_prioritizes_cost_context(self, fetch):
+ source = parse_page(
+ """
+ The AI agent checks each request, updates the CRM, drafts replies, creates follow-up tasks, and notifies the sales team for approval.
+ A simple chatbot costs less than a custom AI agent that connects your website, CRM, email, database, and internal dashboard.
+ Modern AI agents can coordinate tools, memory, retrieval, orchestration, approvals, and connected workflows for growing teams.
+ """,
+ requested_url="https://s.com", final_url="https://s.com", status_code=200,
+ )
+ target = parse_page(
+ """AI Agent Cost in 2026: Pricing Models, Hidden Costs, TCO, and ROI
+ AI Agent Cost in 2026
+ How Much Does an AI Agent Cost?
+ An AI agent can cost less than one thousand dollars per month or require substantial custom development. Pricing, total cost of ownership, operating expense, and ROI depend on integrations, infrastructure, monitoring, and support.
""",
+ requested_url="https://t.com", final_url="https://t.com", status_code=200,
+ )
+ fetch.side_effect = [source, target]
+ items = suggest_placements("https://s.com", "https://t.com", "AI Agent", top_n=3)
+ self.assertEqual(items[0].paragraph_index, 2)
+ self.assertGreater(items[0].destination_score, items[1].destination_score)
+ self.assertIn(items[0].destination_fit, {"high", "very_high"})
+ for item in items[1:]:
+ self.assertEqual(item.destination_fit, "low")
+
-if __name__ == "__main__": unittest.main()
+if __name__ == "__main__":
+ unittest.main()