Claude
Skills
Sign in
Back

neo4j-document-import-skill

Included with Lifetime
$97 forever

Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph. Use when chunking PDFs, HTML, plain text, or Markdown; extracting entities and relationships from text with an LLM (SimpleKGPipeline, neo4j-graphrag); loading JSON via apoc.load.json; building Document→Chunk→Entity graph structures; or connecting LangChain/LlamaIndex document loaders to Neo4j. Covers neo4j-graphrag SimpleKGPipeline, LLM Graph Builder web UI, entity resolution, chunking strategies, and graph schema design for RAG pipelines. Does NOT handle structured CSV/relational import — use neo4j-import-skill. Does NOT handle GraphRAG retrieval after ingestion — use neo4j-graphrag-skill. Does NOT handle vector index creation — use neo4j-vector-search-skill.

Design

What this skill does


# Neo4j Document Import Skill

## When to Use

- Ingesting PDFs, HTML, plain text, Markdown into Neo4j as a knowledge graph
- Chunking documents and storing `:Chunk` nodes with embeddings
- Extracting entities and relationships from text with an LLM
- Using `SimpleKGPipeline` (neo4j-graphrag) programmatically
- Using Neo4j LLM Graph Builder (no-code web UI)
- Loading semi-structured JSON via `apoc.load.json`
- Connecting LangChain or LlamaIndex document loaders to Neo4j

## When NOT to Use

- **Structured CSV / relational data** → `neo4j-import-skill`
- **GraphRAG retrieval after ingestion** → `neo4j-graphrag-skill`
- **Vector index creation** → `neo4j-vector-search-skill`
- **Cypher query writing** → `neo4j-cypher-skill`

---

## Approach Decision Table

| Situation | Approach |
|---|---|
| No code; drag-and-drop UX wanted | LLM Graph Builder web UI |
| Programmatic pipeline; PDFs/text | `SimpleKGPipeline` (neo4j-graphrag) |
| JSON / REST API responses | `apoc.load.json` or Python + UNWIND |
| LangChain already in stack | `Neo4jGraph` + document loader |
| LlamaIndex already in stack | `Neo4jQueryEngine` / `Neo4jVectorStore` |
| Chunk-only (no entity extraction) | Manual chunking + MERGE pattern |

---

## Install

```bash
pip install neo4j-graphrag                   # includes SimpleKGPipeline
pip install neo4j-graphrag[openai]           # + OpenAI LLM/embedder
pip install neo4j-graphrag[anthropic]        # + Anthropic Claude
pip install neo4j-graphrag[google]           # + Vertex AI / Gemini
pip install neo4j-graphrag[bedrock]          # + Amazon Bedrock (boto3) — added v1.15.0
pip install neo4j-graphrag[ollama]           # + Ollama (local)
pip install neo4j-graphrag[mistralai]        # + MistralAI
pip install neo4j-graphrag[fuzzy-matching]   # + FuzzyMatchResolver (rapidfuzz)
# spaCy entity resolver (Python <= 3.13 only — unsupported on 3.14+):
pip install neo4j-graphrag[nlp]
```

Requires: `neo4j>=5.17.0` (driver 6.x supported), Python>=3.10, Neo4j>=5.18.1 (Aura>=5.18.0).

---

## Step 1 — Define Graph Schema

Schema controls what the LLM extracts. Define before pipeline construction.

```python
# Option A — Simple string lists (LLM infers descriptions)
entities = ["Person", "Organization", "Location", "Product", "Event"]
relations = ["WORKS_AT", "LOCATED_IN", "KNOWS", "MENTIONS", "PART_OF"]
patterns = [
    ("Person", "WORKS_AT", "Organization"),
    ("Organization", "LOCATED_IN", "Location"),
    ("Person", "KNOWS", "Person"),
    ("Article", "MENTIONS", "Organization"),
]

# Option B — Rich GraphSchema (production; best extraction quality)
from neo4j_graphrag.experimental.components.schema import (
    GraphSchema, NodeType, RelationshipType, PropertyType, ConstraintType
)
schema = GraphSchema(
    node_types=[
        NodeType(
            label="Person",
            description="A human individual",
            properties=[
                PropertyType(name="name", type="STRING"),
                PropertyType(name="role", type="STRING"),
            ],
        ),
        NodeType(
            label="Organization",
            description="A company or institution",
            properties=[
                PropertyType(name="name", type="STRING"),
                PropertyType(name="industry", type="STRING"),
            ],
        ),
    ],
    relationship_types=[
        RelationshipType(label="WORKS_AT", description="Employment relationship"),
    ],
    patterns=[("Person", "WORKS_AT", "Organization")],
    # Optional: constraints emitted to ParquetWriter metadata (v1.15.0+)
    constraints=[
        ConstraintType(label="Person", property_name="name", type="UNIQUENESS"),
        ConstraintType(label="Organization", property_name="name", type="KEY"),
    ],
)

# Option C — Auto-extract schema from text (no constraints)
schema = "EXTRACTED"   # LLM infers types; noisier output
schema = "FREE"        # No schema guidance; most noise
```

Use Option B for production; Option A for prototyping; `"EXTRACTED"` only for exploration.

---

## Step 2 — SimpleKGPipeline Setup

```python
import asyncio
from neo4j import GraphDatabase
from neo4j_graphrag.experimental.pipeline.kg_builder import SimpleKGPipeline
from neo4j_graphrag.llm import OpenAILLM
from neo4j_graphrag.embeddings import OpenAIEmbeddings

driver = GraphDatabase.driver(
    "neo4j+s://xxxx.databases.neo4j.io",
    auth=("neo4j", "password")
)

llm = OpenAILLM(
    model_name="gpt-4.1",
    model_params={"temperature": 0},
    # Note: SimpleKGPipeline auto-enables structured output for OpenAI/VertexAI LLMs (v1.14.0+)
    # Do NOT set response_format manually — it is managed by the pipeline
)
embedder = OpenAIEmbeddings()   # OPENAI_API_KEY from env

pipeline = SimpleKGPipeline(
    llm=llm,
    driver=driver,
    embedder=embedder,
    schema=schema,              # GraphSchema, dict, "FREE", or "EXTRACTED"
    from_file=True,             # False → pass text= instead of file_path=
    on_error="IGNORE",          # RAISE to surface extraction failures
    perform_entity_resolution=True,
    neo4j_database="neo4j",     # omit to use default
)
```

**LLM alternatives** (same interface):
- `AnthropicLLM(model_name="claude-3-5-sonnet-20241022")`
- `VertexAILLM(model_name="gemini-2.0-flash")`
- `OllamaLLM(model_name="llama3")` — local; no API key needed
- `BedrockLLM(model_id="anthropic.claude-3-5-sonnet-20241022-v2:0")` — Amazon Bedrock (v1.15.0+)

---

## Step 3 — Run the Pipeline

```python
# From PDF file:
result = asyncio.run(pipeline.run_async(
    file_path="report.pdf",        # auto-dispatches to PdfLoader
    document_metadata={"source": "Q4 report", "year": 2025},
))

# From Markdown file (v1.15.0+):
result = asyncio.run(pipeline.run_async(
    file_path="notes.md",          # auto-dispatches to MarkdownLoader
    document_metadata={"source": "meeting notes"},
))

# Note: old `from_pdf=True` parameter is DEPRECATED since v1.15.0; use `from_file=True` instead
# pipeline = SimpleKGPipeline(..., from_file=True)   ← correct
# pipeline = SimpleKGPipeline(..., from_pdf=True)    ← deprecated

# From raw text:
result = asyncio.run(pipeline.run_async(
    text=document_text,
))

# Batch — process multiple files:
async def ingest_all(paths):
    for p in paths:
        await pipeline.run_async(file_path=str(p))

asyncio.run(ingest_all(list(pdf_dir.glob("*.pdf"))))
```

`document_metadata` dict is stored as properties on the `:Document` node.

---

## Step 4 — Chunking Configuration

Default splitter: `FixedSizeSplitter(chunk_size=300, chunk_overlap=50)`.

```python
from neo4j_graphrag.experimental.components.text_splitters.fixed_size_splitter import FixedSizeSplitter

splitter = FixedSizeSplitter(
    chunk_size=512,       # tokens; 300–512 typical for GPT-4o
    chunk_overlap=50,     # ~10% of chunk_size; preserves boundary context
    approximate=True,     # respect sentence/word boundaries when possible
)

pipeline = SimpleKGPipeline(
    ...,
    text_splitter=splitter,
)
```

Chunking guidance:
| Document type | chunk_size | chunk_overlap |
|---|---|---|
| Dense technical text | 256–512 | 50–80 |
| Narrative / news articles | 512–1024 | 80–128 |
| Legal / financial docs | 256–384 | 40–64 |

Rule: chunk must fit within LLM context for extraction + within embedding model limits. GPT-4o: 128k context; `text-embedding-3-small`: 8191 tokens. Never set chunk_size > 2048.

---

## Step 5 — Entity Resolution

Merge duplicate extracted entities after pipeline run.

```python
from neo4j_graphrag.experimental.components.resolver import (
    SinglePropertyExactMatchResolver,   # identical name → merge
    FuzzyMatchResolver,                  # Levenshtein similarity; needs rapidfuzz
    SpaCySemanticMatchResolver,          # cosine similarity; needs neo4j-graphrag[nlp]
)

# Exact match (fastest; good baseline)
resolver = SinglePropertyExactMatchResolver(driver)
asyncio.run(resolver.run())

# Fuzzy match (handles typos / alternate spellings)
from neo4j_graphrag.experimental.components.resolver i

Related in Design