{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "\"Open" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "# AnalyticDB\n", "\n", ">[AnalyticDB for PostgreSQL](https://www.alibabacloud.com/help/en/analyticdb-for-postgresql/product-overview/overview-product-overview) is a massively parallel processing (MPP) data warehousing service that is designed to analyze large volumes of data online.\n", "\n", "\n", "To run this notebook you need a AnalyticDB for PostgreSQL instance running in the cloud (you can get one at [common-buy.aliyun.com](https://common-buy.aliyun.com/?commodityCode=GreenplumPost®ionId=cn-hangzhou&request=%7B%22instance_rs_type%22%3A%22ecs%22%2C%22engine_version%22%3A%226.0%22%2C%22seg_node_num%22%3A%224%22%2C%22SampleData%22%3A%22false%22%2C%22vector_optimizor%22%3A%22Y%22%7D)).\n", "\n", "After creating the instance, you should create a manager account by [API](https://www.alibabacloud.com/help/en/analyticdb-for-postgresql/developer-reference/api-gpdb-2016-05-03-createaccount) or 'Account Management' at the instance detail web page.\n", "\n", "You should ensure you have `llama-index` installed:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "%pip install llama-index-vector-stores-analyticdb" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "!pip install llama-index" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Please provide parameters:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "import os\n", "import getpass\n", "\n", "# alibaba cloud ram ak and sk:\n", "alibaba_cloud_ak = \"\"\n", "alibaba_cloud_sk = \"\"\n", "\n", "# instance information:\n", "region_id = \"cn-hangzhou\" # region id of the specific instance\n", "instance_id = \"gp-xxxx\" # adb instance id\n", "account = \"test_account\" # instance account name created by API or 'Account Management' at the instance detail web page\n", "account_password = \"\" # instance account password" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Import needed package dependencies:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from llama_index.core import (\n", " VectorStoreIndex,\n", " SimpleDirectoryReader,\n", " StorageContext,\n", ")\n", "from llama_index.vector_stores.analyticdb import AnalyticDBVectorStore" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Load some example data:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "!mkdir -p 'data/paul_graham/'\n", "!wget 'https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/paul_graham/paul_graham_essay.txt' -O 'data/paul_graham/paul_graham_essay.txt'" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Read the data:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# load documents\n", "documents = SimpleDirectoryReader(\"./data/paul_graham/\").load_data()\n", "print(f\"Total documents: {len(documents)}\")\n", "print(f\"First document, id: {documents[0].doc_id}\")\n", "print(f\"First document, hash: {documents[0].hash}\")\n", "print(\n", " \"First document, text\"\n", " f\" ({len(documents[0].text)} characters):\\n{'='*20}\\n{documents[0].text[:360]} ...\"\n", ")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Create the AnalyticDB Vector Store object:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "analytic_db_store = AnalyticDBVectorStore.from_params(\n", " access_key_id=alibaba_cloud_ak,\n", " access_key_secret=alibaba_cloud_sk,\n", " region_id=region_id,\n", " instance_id=instance_id,\n", " account=account,\n", " account_password=account_password,\n", " namespace=\"llama\",\n", " collection=\"llama\",\n", " metrics=\"cosine\",\n", " embedding_dimension=1536,\n", ")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Build the Index from the Documents:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "storage_context = StorageContext.from_defaults(vector_store=analytic_db_store)\n", "\n", "index = VectorStoreIndex.from_documents(\n", " documents, storage_context=storage_context\n", ")" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Query using the index:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "query_engine = index.as_query_engine()\n", "response = query_engine.query(\"Why did the author choose to work on AI?\")\n", "\n", "print(response.response)" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Delete the collection:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "analytic_db_store.delete_collection()" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3 (ipykernel)", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3" } }, "nbformat": 4, "nbformat_minor": 4 }