documentation update
This commit is contained in:
@@ -99,7 +99,7 @@ returns without writing anything.
|
||||
|
||||
For each entity in `parsed.entities`:
|
||||
|
||||
1. Validate `name`, `type` (must be in `ENTITY_TYPES`), and not in `IGNORED_NAMES`
|
||||
1. Validate `name`, `type` (must be in `ENTITY_TYPES`), and not a greeting via `isIgnoredName(name)`
|
||||
2. Call `upsertEntity(name, type, notes)`:
|
||||
- **Insert**: creates new row with `mention_count = 1`, `source = 'extraction'`
|
||||
- **Conflict** on `(name, type)`: increments `mention_count`, updates `last_seen_at`, preserves existing `notes` if new extraction returns null
|
||||
@@ -109,7 +109,27 @@ For each entity in `parsed.entities`:
|
||||
|
||||
**Valid entity types:** `person`, `place`, `project`, `technology`, `concept`, `organization`
|
||||
|
||||
**Stoplist (ignored names):** `good morning`, `good night`, `hello`, `goodbye`, `thanks`, `thank you`
|
||||
**Greeting filter:** `isIgnoredName(name)` normalizes the candidate before
|
||||
matching — lowercased, punctuation stripped (`/[^\w\s]/g`), trimmed — then
|
||||
tests membership against the `IGNORED_NAMES` set. Normalizing means variants
|
||||
like `"Good morning!"` or `" hello "` are caught, not just exact lowercase
|
||||
matches. Current set: `good morning`, `good night`, `good evening`,
|
||||
`good afternoon`, `hello`, `hi`, `hey`, `goodbye`, `bye`, `thanks`,
|
||||
`thank you`, `morning`.
|
||||
|
||||
The extraction prompt also instructs the model not to emit greetings or
|
||||
conversational filler as entities, so the filter is a backstop rather than the
|
||||
sole defense — necessary because the model (qwen2.5:3b) will occasionally label
|
||||
a greeting as a `concept` or `topic`, both valid types that the `ENTITY_TYPES`
|
||||
check alone won't reject.
|
||||
|
||||
> **Note:** the filter runs only at extraction time; it does not retroactively
|
||||
> remove greetings stored before the filter (or an entry) existed. Pre-filter
|
||||
> stragglers must be deleted manually via `DELETE /entities/:id`, which now
|
||||
> also removes the Qdrant vector (see `memory-service.md` → Delete Behaviour).
|
||||
> `\w` is ASCII-only, so non-ASCII greetings (e.g. `buenos días`) normalize
|
||||
> imperfectly — switch to `/[^\p{L}\p{N}\s]/gu` if the set gains non-ASCII
|
||||
> entries.
|
||||
|
||||
## Relationship Processing
|
||||
|
||||
@@ -137,4 +157,4 @@ All steps after the initial model call are wrapped in a single outer try/catch.
|
||||
If Ollama is unreachable, returns a non-200 status, or the JSON cannot be
|
||||
parsed, the function logs at `warn` level and returns. There is no retry logic.
|
||||
Individual entity embedding failures are caught per-entity and logged at `warn`
|
||||
level without affecting other entities in the same batch.
|
||||
level without affecting other entities in the same batch.
|
||||
Reference in New Issue
Block a user