chore: import upstream snapshot with attribution
Integ / changes (push) Has been skipped
Pre-commit / pre-commit (push) Failing after 1s
CLI exit codes / changes (push) Has been skipped
Test (Install) / changes (push) Has been skipped
Test (Python) / changes (push) Has been skipped
Test (TypeScript) / changes (push) Has been skipped
CLI exit codes / cli-gate (push) Has been cancelled
Test (Install) / test-install-gate (push) Has been cancelled
Integ / integ-gate (push) Has been cancelled
Test (Python) / test-python-gate (push) Has been cancelled
Test (TypeScript) / test-typescript-gate (push) Has been cancelled
Test (Install) / python-minimal (3.12) (push) Has been cancelled
Test (Install) / python-minimal (3.11) (push) Has been cancelled
Test (Install) / python-extra (agno, mirage.agents.agno) (push) Has been cancelled
Test (Install) / python-extra (chroma, mirage.resource.chroma) (push) Has been cancelled
Test (Install) / python-extra (pdf, mirage.core.filetype.pdf) (push) Has been cancelled
Integ / integ (push) Has been cancelled
Integ / integ-database (push) Has been cancelled
Integ / integ-database-ts (push) Has been cancelled
Integ / integ-data (push) Has been cancelled
Integ / integ-ssh (push) Has been cancelled
Integ / integ-ssh-ts (push) Has been cancelled
Test (Python) / audit (push) Has been cancelled
Test (TypeScript) / test (push) Has been cancelled
Test (TypeScript) / python-fs-shim (push) Has been cancelled
CLI exit codes / Python CLI (push) Has been cancelled
CLI exit codes / TypeScript CLI (push) Has been cancelled
CLI exit codes / Cross-language snapshot interop (push) Has been cancelled
Test (Python) / test (push) Has been cancelled
Test (Python) / import-isolation (deepagents, openai, mirage.agents.openai_agents) (push) Has been cancelled
Test (Python) / import-isolation (deepagents, pydantic-ai, mirage.agents.pydantic_ai) (push) Has been cancelled
Integ / integ-ts (push) Has been cancelled
Integ / integ-fuse (push) Has been cancelled
Test (Install) / python-extra (databricks, mirage.resource.databricks_volume) (push) Has been cancelled
Test (Install) / python-extra (deepagents, mirage.agents.langchain) (push) Has been cancelled
Test (Install) / python-extra (email, mirage.resource.email) (push) Has been cancelled
Test (Install) / python-extra (fuse, mirage.fuse.mount) (push) Has been cancelled
Test (Install) / python-extra (hdf5, mirage.core.filetype.hdf5) (push) Has been cancelled
Test (Install) / python-extra (hf, mirage.resource.hf_buckets) (push) Has been cancelled
Test (Install) / python-extra (lancedb, mirage.resource.lancedb) (push) Has been cancelled
Test (Install) / python-extra (langfuse, mirage.resource.langfuse) (push) Has been cancelled
Test (Install) / python-extra (mongodb, mirage.resource.mongodb) (push) Has been cancelled
Test (Install) / python-extra (nextcloud, mirage.resource.nextcloud) (push) Has been cancelled
Test (Install) / python-extra (openai, mirage.agents.openai_agents) (push) Has been cancelled
Test (Install) / python-extra (openhands, mirage.agents.openhands, 3.12) (push) Has been cancelled
Test (Install) / python-extra (parquet, mirage.core.filetype.parquet) (push) Has been cancelled
Test (Install) / python-extra (postgres, mirage.resource.postgres) (push) Has been cancelled
Test (Install) / python-extra (pydantic-ai, mirage.agents.pydantic_ai) (push) Has been cancelled
Test (Install) / python-extra (qdrant, mirage.resource.qdrant) (push) Has been cancelled
Test (Install) / python-extra (redis, mirage.resource.redis) (push) Has been cancelled
Test (Install) / python-extra (s3, mirage.resource.s3) (push) Has been cancelled
Test (Install) / python-extra (ssh, mirage.resource.ssh) (push) Has been cancelled
Test (Install) / ts-minimal (push) Has been cancelled

This commit is contained in:
wehub-resource-sync
2026-07-13 12:30:44 +08:00
commit bcbd1bdb22
5748 changed files with 562488 additions and 0 deletions
+13
View File
@@ -0,0 +1,13 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
+29
View File
@@ -0,0 +1,29 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
from mirage.accessor.lancedb import LanceDBAccessor
from mirage.cache.index import IndexCacheStore
from mirage.commands.builtin.constants import SCOPE_ERROR
from mirage.core.lancedb.readdir import readdir
from mirage.types import PathSpec
from mirage.utils.glob_walk import resolve_glob_with
async def resolve_glob(
accessor: LanceDBAccessor,
paths: list[PathSpec],
index: IndexCacheStore | None = None,
) -> list[PathSpec]:
return await resolve_glob_with(readdir, accessor, paths, index,
SCOPE_ERROR)
+82
View File
@@ -0,0 +1,82 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
from mirage.accessor.lancedb import LanceDBAccessor
def _quote(value: str) -> str:
return value.replace("'", "''")
def _eq(column: str, value: str) -> str:
text = str(value)
if text.lstrip("-").isdigit():
return f"{column} = {text}"
return f"{column} = '{_quote(text)}'"
def _where(filters: dict[str, str]) -> str:
return " AND ".join(_eq(col, val) for col, val in filters.items())
async def list_tables(accessor: LanceDBAccessor) -> list[str]:
db = await accessor.db()
result = await db.list_tables()
names = result.tables if hasattr(result, "tables") else result
return sorted(names)
async def table_exists(accessor: LanceDBAccessor, name: str) -> bool:
return name in await list_tables(accessor)
async def distinct_values(accessor: LanceDBAccessor, table: str, column: str,
filters: dict[str, str], limit: int) -> list[str]:
tbl = await accessor.table(table)
query = tbl.query().select([column]).limit(limit)
if filters:
query = query.where(_where(filters))
rows = await query.to_list()
values = {str(row[column]) for row in rows if row.get(column) is not None}
return sorted(values)
async def rows_matching(accessor: LanceDBAccessor, table: str,
filters: dict[str, str], columns: list[str],
limit: int) -> list[dict]:
tbl = await accessor.table(table)
query = tbl.query().select(columns).limit(limit)
if filters:
query = query.where(_where(filters))
return await query.to_list()
async def row_record(accessor: LanceDBAccessor, table: str, id_column: str,
row_id: str) -> dict | None:
tbl = await accessor.table(table)
rows = await tbl.query().where(_eq(id_column, row_id)).limit(1).to_list()
return rows[0] if rows else None
async def search_rows(accessor: LanceDBAccessor, table: str, query_text: str,
limit: int) -> list[dict]:
key = (table, query_text, limit)
cached = accessor.cached_search(key)
if cached is not None:
return cached
tbl = await accessor.table(table)
builder = await tbl.search(query_text)
rows = await builder.limit(limit).to_list()
accessor.store_search(key, rows)
return rows
+61
View File
@@ -0,0 +1,61 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
import base64
from mirage.accessor.lancedb import LanceDBAccessor
from mirage.cache.index import IndexCacheStore
from mirage.core.lancedb.query import row_record
from mirage.core.lancedb.render import render_card
from mirage.core.lancedb.scope import ScopeLevel, detect_scope
from mirage.types import PathSpec
from mirage.utils.errors import enoent
async def _resolve_row(accessor: LanceDBAccessor, scope, config,
virtual: str) -> dict:
row = await row_record(accessor, scope.table, config.id_column,
scope.row_id)
if row is None:
raise enoent(virtual)
return row
def _blob_bytes(value: object) -> bytes:
if isinstance(value, bytes):
return value
if isinstance(value, str):
return base64.b64decode(value)
raise ValueError("blob column is not bytes or base64 str")
async def read(
accessor: LanceDBAccessor,
path: PathSpec,
index: IndexCacheStore = None,
) -> bytes:
if isinstance(path, str):
path = PathSpec(virtual=path,
directory=path,
resource_path=path.strip("/"))
config = accessor.config
scope = detect_scope(path, config)
if scope.level != ScopeLevel.ROW:
raise enoent(path)
row = await _resolve_row(accessor, scope, config, path.virtual)
if scope.blob:
if not config.blob_column:
raise enoent(path)
return _blob_bytes(row.get(config.blob_column))
return render_card(row, config)
+74
View File
@@ -0,0 +1,74 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
from mirage.accessor.lancedb import LanceDBAccessor
from mirage.cache.index import IndexCacheStore
from mirage.core.lancedb.query import (distinct_values, list_tables,
rows_matching)
from mirage.core.lancedb.scope import ScopeLevel, detect_scope
from mirage.types import PathSpec
def is_dir_name(child: str, config) -> bool:
# Row files are recognized by extension, so classification never needs
# the stat fallback.
name = child.rsplit("/", 1)[-1]
if name.endswith(".md"):
return False
if config.blob_column and name.endswith("." + config.blob_ext):
return False
return True
def _row_files(rows: list[dict], config) -> list[str]:
names: list[str] = []
for row in rows:
rid = row[config.id_column]
names.append(f"{rid}.md")
if config.blob_column:
names.append(f"{rid}.{config.blob_ext}")
return names
async def readdir(
accessor: LanceDBAccessor,
path: PathSpec,
index: IndexCacheStore = None,
) -> list[str]:
if isinstance(path, str):
path = PathSpec(virtual=path,
directory=path,
resource_path=path.strip("/"))
config = accessor.config
scope = detect_scope(path, config)
base = path.virtual.rstrip("/")
if scope.level == ScopeLevel.ROOT:
names = await list_tables(accessor)
return [f"{base}/{name}" for name in names]
if scope.level == ScopeLevel.GROUP_DIR:
depth = len(scope.filters)
total = len(config.group_by)
if depth < total:
names = await distinct_values(accessor, scope.table,
config.group_by[depth],
scope.filters, config.max_rows)
else:
rows = await rows_matching(accessor, scope.table, scope.filters,
[config.id_column], config.max_rows)
names = _row_files(rows, config)
return [f"{base}/{name}" for name in names]
raise FileNotFoundError(path.virtual)
+37
View File
@@ -0,0 +1,37 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
from mirage.resource.lancedb.config import LanceDBConfig
_SKIP_KEYS = {"_distance", "_rowid", "_score"}
def render_card(row: dict, config: LanceDBConfig) -> bytes:
lines: list[str] = []
title = row.get(config.title_column) if config.title_column else None
if title is not None:
lines.append(f"# {title}")
lines.append("")
for key, value in row.items():
if key in _SKIP_KEYS:
continue
if key == config.vector_column or key == config.blob_column:
continue
lines.append(f"{key}: {value}")
if config.blob_column and config.id_column in row:
lines.append(f"blob: {row[config.id_column]}.{config.blob_ext}")
distance = row.get("_distance")
if distance is not None:
lines.append(f"score: {float(distance):.4f}")
return ("\n".join(lines) + "\n").encode()
+85
View File
@@ -0,0 +1,85 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
from dataclasses import dataclass, field
from enum import Enum
from mirage.resource.lancedb.config import LanceDBConfig
from mirage.types import PathSpec
class ScopeLevel(str, Enum):
ROOT = "root"
GROUP_DIR = "group_dir"
ROW = "row"
UNKNOWN = "unknown"
@dataclass
class LanceDBScope:
level: ScopeLevel
table: str | None = None
filters: dict[str, str] = field(default_factory=dict)
row_id: str | None = None
blob: bool = False
resource_path: str = "/"
def _parse_row_file(name: str,
config: LanceDBConfig) -> tuple[str, bool] | None:
if name.endswith(".md"):
return name[:-len(".md")], False
if config.blob_column:
suffix = "." + config.blob_ext
if name.endswith(suffix):
return name[:-len(suffix)], True
return None
def detect_scope(path, config: LanceDBConfig) -> LanceDBScope:
raw = path.mount_path if isinstance(path, PathSpec) else path
key = raw.strip("/")
segs = key.split("/") if key else []
if config.table:
table = config.table
rest = segs
else:
if not segs:
return LanceDBScope(level=ScopeLevel.ROOT, resource_path=raw)
table = segs[0]
rest = segs[1:]
gb = config.group_by
n = len(gb)
if len(rest) <= n:
filters = {gb[i]: rest[i] for i in range(len(rest))}
return LanceDBScope(level=ScopeLevel.GROUP_DIR,
table=table,
filters=filters,
resource_path=raw)
if len(rest) == n + 1:
filters = {gb[i]: rest[i] for i in range(n)}
parsed = _parse_row_file(rest[n], config)
if parsed is not None:
return LanceDBScope(level=ScopeLevel.ROW,
table=table,
filters=filters,
row_id=parsed[0],
blob=parsed[1],
resource_path=raw)
return LanceDBScope(level=ScopeLevel.UNKNOWN, resource_path=raw)
+81
View File
@@ -0,0 +1,81 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
from mirage.accessor.lancedb import LanceDBAccessor
from mirage.core.lancedb.query import search_rows
from mirage.core.lancedb.render import render_card
from mirage.resource.lancedb.config import LanceDBConfig
from mirage.types import PathSpec
def _target_table(paths: list[PathSpec], config: LanceDBConfig) -> str | None:
if config.table:
return config.table
for path in paths:
raw = path.mount_path if isinstance(path, PathSpec) else path
key = raw.strip("/")
if key:
return key.split("/")[0]
return None
def _canonical_path(row: dict, config: LanceDBConfig, table: str,
mount_prefix: str) -> str:
segs: list[str] = []
if not config.table:
segs.append(str(table))
for column in config.group_by:
if column in row and row[column] is not None:
segs.append(str(row[column]))
segs.append(f"{row[config.id_column]}.md")
prefix = mount_prefix.rstrip("/")
return prefix + "/" + "/".join(segs)
def _block(row: dict, config: LanceDBConfig, table: str,
mount_prefix: str) -> str:
path = _canonical_path(row, config, table, mount_prefix)
distance = row.get("_distance")
header = path if distance is None else f"{path}:{float(distance):.4f}"
body_row = {k: v for k, v in row.items() if k != "_distance"}
content = render_card(body_row, config).decode().rstrip("\n")
return f"{header}\n{content}"
async def search_rows_output(
accessor: LanceDBAccessor,
query: str,
paths: list[PathSpec],
top_k: int,
threshold: float,
mount_prefix: str,
) -> bytes:
if not query:
raise ValueError("search: query is required")
if top_k <= 0:
raise ValueError("search: top-k must be positive")
table = _target_table(paths, accessor.config)
if table is None:
raise FileNotFoundError("search: no table to search")
rows = await search_rows(accessor, table, query, top_k)
blocks: list[str] = []
for row in rows:
distance = row.get("_distance")
if threshold > 0 and distance is not None and float(
distance) > threshold:
continue
blocks.append(_block(row, accessor.config, table, mount_prefix))
if not blocks:
return b""
return ("\n".join(blocks) + "\n").encode()
+61
View File
@@ -0,0 +1,61 @@
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
from mirage.accessor.lancedb import LanceDBAccessor
from mirage.cache.index import IndexCacheStore
from mirage.core.lancedb.query import table_exists
from mirage.core.lancedb.read import read
from mirage.core.lancedb.scope import ScopeLevel, detect_scope
from mirage.types import FileStat, FileType, PathSpec
_IMAGE_TYPES = {
"png": FileType.IMAGE_PNG,
"jpg": FileType.IMAGE_JPEG,
"jpeg": FileType.IMAGE_JPEG,
"gif": FileType.IMAGE_GIF,
}
def _name_of(path: PathSpec) -> str:
stripped = path.virtual.rstrip("/")
return stripped.rsplit("/", 1)[-1] or "/"
async def stat(
accessor: LanceDBAccessor,
path: PathSpec,
index: IndexCacheStore = None,
) -> FileStat:
if isinstance(path, str):
path = PathSpec(virtual=path,
directory=path,
resource_path=path.strip("/"))
config = accessor.config
scope = detect_scope(path, config)
if scope.level == ScopeLevel.UNKNOWN:
raise FileNotFoundError(path.virtual)
if scope.table and not await table_exists(accessor, scope.table):
raise FileNotFoundError(path.virtual)
if scope.level in (ScopeLevel.ROOT, ScopeLevel.GROUP_DIR):
return FileStat(name=_name_of(path), type=FileType.DIRECTORY)
data = await read(accessor, path, index)
if scope.blob:
file_type = _IMAGE_TYPES.get(config.blob_ext, FileType.BINARY)
else:
file_type = FileType.TEXT
return FileStat(name=_name_of(path), size=len(data), type=file_type)